Warm tip: This article is reproduced from serverfault.com, please click

grep sed ubuntu

Find multiple occurences of same string

发布于 2020-11-29 21:13:35

In a range of files, I want to see which line has atleast 4 times the same occurence of the same word. This word can be any word.

So input:

a a a b b e e e
o o o o p p p y y y
w r r r u u i i o o r
x x o o i i p p z z y y

Output:

o o o o p p p y y y
w r r r u u i i o o r

What I have tried at the moment is to make sure that sentences are put separate, ready to be processed basically.

cat * |
    tr '\n' ' '|
    sed 's/[.!?;"]/ & /g' |
    sed 's/[.!?]/&\n/g'|
    grep -E -w '\b([[:alnum:]]*)\{4*\}\b'

But my grep doesn't get anything, so how do I get that Grep only prints out all sentences which contain a word which occurs atleast 4 times in it?

Questioner

Hooiberg12

Viewed

0

Wiktor Stribiżew 2020-11-30 06:17:22

With GNU grep, you can use a PCRE regex like

grep -P '\b(\w+)\b(.*\b\1\b){3}'

See the regex demo.

Test in Ubuntu 18.04.4 LTS:

Details

\b(\w+)\b - a whole word (captured in Group 1) (\b is a word boundary and \w matches letters, digits or underscores)
(.*\b\1\b){3} - three occurrences ({3}) of any text followed with the same value as in Group 1 (as \1 is an inline backreference to Group 1 value) as a whole word (again, \b word boundaries are used.)

wjandrea 2020-11-29 21:30:44

You can simplify by putting the word bounds in the group: grep -E '(\b\w\b)(.*\1){3}'. Unless there's an edge case I haven't thought of.

Wiktor Stribiżew 2020-11-29 21:32:34

@wjandrea The capturing group only keeps the value, not the pattern. So, \1 is unaware of the fact if the string it holds was captured as a whole word or not. We need all the word boundaries I used in the pattern.

wjandrea 2020-11-29 21:35:59

Ah, I see, if you use input like o do to moe, mine matches it (false positive), yours doesn't.

Hooiberg12 2020-11-29 22:11:39

I am also getting false positives with yours solution @WiktorStribiżew. It most likely has to do with the first word boundary, I would say. en vrouwen gelijk voor de wet en maken we geen is one of my results, but the word en does not pop up separately 4 times. It does if you count the en inside of the words.

Hooiberg12 2020-11-29 22:20:53

Ah yes this seems to work. Thank you for the solution to my problem.

热门帖子

1

推荐一些好玩的/大众的手游

2

求指教后端项目迁移方案

3

迷你洗衣机是不是都是智商税？

4

求助一个排查了半年没解决的 MySQL order by 子句导致索引失效的问题， 500 多万条记录的小表要查快两分钟

5

个人开发了一款 WordPress 主题： iPao，集成了 AI 总结功能

6

偶然发现奇游加速器会在系统里植入根证书

7

国内有蒲公英替代品推荐吗？

8

语音助手这个东西真的会监听谈话并且上传，从而泄漏隐私吗？

9

出一些有意思的域名-明盘

10

jetbrains 全家桶升级 2024 后，在滚动代码时候感觉有点掉帧

热门github

1

A multi-platform library for OpenGL, OpenGL ES, Vulkan, window and input

2

Dev tool that writes scalable apps from scratch while the developer oversees the implementation

3

shadcn/ui, but for Svelte. ✨

4

The Python Risk Identification Tool for generative AI (PyRIT) is an open access automation framework to empower security professionals and machine learning engineers to proactively find risks in their generative AI systems.

5

Performance-portable, length-agnostic SIMD with runtime dispatch

6

ZK Credo

7

OpenCodeInterpreter: Integrating Code Generation with Execution and Refinement

8

Joplin - the secure note taking and to-do app with synchronisation capabilities for Windows, macOS, Linux, Android and iOS.

9

Mamba is a new state space model architecture showing promising performance on information-dense data such as language modeling, where previous subquadratic models fall short of Transformers. It is based on the line of progress on structured state space models, with an efficient hardware-aware design and implementation in the spirit of FlashAttention.

10

This repository contains System Design resources which are useful while preparing for interviews and learning Distributed Systems

11

Curso para aprender el lenguaje de programación Python desde cero y para principiantes. 75 clases, 37 horas en vídeo, código, proyectos y grupo de chat. Fundamentos, frontend, backend, testing, IA...

12

🎓 Path to a free self-taught education in Computer Science!

13

1️⃣🐝🏎️ The One Billion Row Challenge -- A fun exploration of how quickly 1B rows from a text file can be aggregated with Java

14

A collective list of free APIs

15

📚 Freely available programming books