---
tags:
- cli
- l1
- flashcard-deck
- regex
---
<!-- wiki:breadcrumb:start -->
[Portal](../../../../library/portal/index.md) | **Level:** [L1: Foundations](../../../../library/portal/levels.md) | **Topics:** [Regex & Text Wrangling](../../../../library/portal/topics.md) | **Domain:** CLI Tools
<!-- wiki:breadcrumb:end -->

id	category	difficulty	tags	question	answer	source_path
regex/dbb48184b176	regex	easy	regex, grep, basics	What grep flag enables Extended Regular Expressions (ERE) to avoid backslash escaping?	grep -E (or egrep). ERE allows +, ?, {n,m}, and () without backslash escaping, unlike BRE where you must write \+, \?, \{n,m\}, and \(\).\n\nRemember: "Regex = pattern matching language." The dot . matches any character except newline (unless DOTALL flag).\n\nExample: grep -E "^error" log.txt matches lines starting with "error."	training/library/topics/regex-text-wrangling/primer.md
regex/cc35296346a7	regex	easy	regex, sed, substitution	How do you replace all occurrences of a pattern on every line in a file using sed?	sed 's/old/new/g' file. The g flag means global (all occurrences per line). Without g, only the first occurrence on each line is replaced.\n\nRemember: "Star = zero or more, Plus = one or more, Question = zero or one."\n\nExample: colou?r matches both "color" and "colour". The ? makes the u optional.	training/library/topics/regex-text-wrangling/primer.md
regex/0cd3962c1bfd	regex	easy	regex, awk, columns	How do you print the second column of a space-delimited file using awk?	awk '{print $2}' file. Use -F to change the delimiter, e.g., awk -F: '{print $1, $3}' /etc/passwd for colon-delimited files.\n\nRemember: "Brackets = character class = pick one." [aeiou] matches any vowel. [^aeiou] matches any non-vowel.\n\nExample: [0-9]{3}-[0-9]{4} matches phone-number-like patterns like 555-1234.	training/library/topics/regex-text-wrangling/primer.md
regex/c62a77738d35	regex	medium	regex, dialects, differences	What are the three regex dialects and their key escaping differences?	BRE (Basic, used by grep/sed): grouping \(\), quantifier \+, \?. ERE (Extended, used by grep -E, awk): grouping (), quantifier +, ?. PCRE (Perl, used by grep -P): same as ERE plus lookahead (?=...), (?!...), non-greedy *?, +?, and \d, \w, \s shortcuts.\n\nGotcha: .* is greedy by default — it matches as much as possible. Use .*? for non-greedy (lazy) matching.\n\nRemember: "Greedy grabs all, lazy stops at first match."	training/library/topics/regex-text-wrangling/primer.md
regex/09772990c40b	regex	medium	regex, sed, capture-groups	How do you use capture groups and backreferences in sed to reformat a date?	echo "2024-03-15" | sed -E 's/([0-9]{4})-([0-9]{2})-([0-9]{2})/\3\/\2\/\1/' outputs 15/03/2024. Parentheses capture groups and \1, \2, \3 reference them. Use -E for extended regex to avoid escaping parentheses.\n\nRemember: "Parens capture, pipe alternates." (cat|dog) matches "cat" or "dog" and captures the match.	training/library/topics/regex-text-wrangling/primer.md
regex/4e4f82fa7107	regex	medium	regex, awk, aggregation	How do you count occurrences by key and compute a sum using awk?	Count by key: awk '{count[$1]++} END {for (k in count) print k, count[k]}'. Sum a column: awk '{sum += $3} END {print sum}'. Average: awk '{sum += $3; n++} END {print sum/n}'. Awk processes line by line, running the main block per line and END once after all input.	training/library/topics/regex-text-wrangling/primer.md
regex/9c757b061477	regex	medium	regex, grep, pipeline	How do you extract all IP addresses from a log file, count unique occurrences, and show the top 20?	grep -oE '[0-9]+\.[0-9]+\.[0-9]+\.[0-9]+' access.log | sort | uniq -c | sort -rn | head -20. The -o flag outputs only matching text, one per line. sort | uniq -c counts unique lines, and sort -rn orders by count descending.	training/library/topics/regex-text-wrangling/primer.md
regex/af496a22598c	regex	hard	regex, posix, portability	What are POSIX character classes and why should you use them over shortcuts like \d?	POSIX classes like [[:digit:]], [[:alpha:]], [[:alnum:]], [[:space:]] are portable across sed, awk, and grep in all locales. The \d, \w, \s shortcuts are PCRE-only (grep -P) and not available in standard sed or awk. For portable scripts, always use POSIX classes.	training/library/topics/regex-text-wrangling/primer.md
regex/911d0a9b0d7a	regex	hard	regex, awk, log-analysis	How would you use awk to find slow requests in an Apache log and compute average response time?	Slow requests (> 1 second, assuming time in column 10): awk '$10 > 1000 {print $7, $10"ms"}' access.log. Average response time: awk '{sum += $10; n++} END {printf "avg: %.2fms\n", sum/n}' access.log. Status code distribution: awk '{print $9}' access.log | sort | uniq -c | sort -rn.	training/library/topics/regex-text-wrangling/primer.md
regex/9891ce72b3d7	regex	hard	regex, sed, advanced	How do you perform operations on a range of lines between two patterns using sed?	sed '/START/,/END/s/foo/bar/g' file replaces foo with bar only on lines between START and END patterns (inclusive). You can also delete a range: sed '/START/,/END/d'. Insert before a match: sed '/pattern/i text'. Append after: sed '/pattern/a text'. Combine with -n and p to extract ranges: sed -n '/START/,/END/p'.	training/library/topics/regex-text-wrangling/primer.md
regex/a1b2c3d4e5f7	regex	easy	regex, character-classes, basics	What are the most common regex character classes and what do they match?	. matches any character (except newline). \d matches a digit [0-9]. \w matches a word character [a-zA-Z0-9_]. \s matches whitespace (space, tab, newline). \b matches a word boundary. Capitalize to negate: \D = non-digit, \W = non-word, \S = non-whitespace. \nNote: \d, \w, \s are PCRE-only — use [[:digit:]], [[:alnum:]], [[:space:]] in POSIX tools like sed and awk.	
regex/b2c3d4e5f6a8	regex	medium	regex, quantifiers, greedy-lazy	What is the difference between greedy and lazy quantifiers in regex?	Greedy quantifiers (*, +, {n,m}) match as much as possible. Lazy quantifiers (*?, +?, {n,m}?) match as little as possible. \nExample: given "bold", the greedy pattern <.*> matches "bold" (entire string), while <.*?> matches "<b>" (first tag only). Lazy quantifiers require PCRE (grep -P). In BRE/ERE, use negated character classes instead: <[^>]*> matches a single tag.	
regex/c3d4e5f6a7b9	regex	medium	regex, capture-groups, backreferences	How do capture groups and backreferences work in regex?	Parentheses () create capture groups that save matched text. Backreferences (\1, \2) reuse captured text. \nExample: (.)\1 matches any character followed by itself (aa, bb, 11). In sed: echo "John Smith" | sed -E 's/(\w+) (\w+)/\2, \1/' outputs "Smith, John". Non-capturing groups (?:...) group without saving — available in PCRE only.	
regex/d4e5f6a7b8c0	regex	hard	regex, lookahead, lookbehind	What are lookahead and lookbehind assertions and when are they useful?	Lookahead (?=...) matches a position followed by a pattern without consuming it. Lookbehind (?<=...) matches preceded by a pattern. Negative versions: (?!...) and (?<!...). \nExample: \d+(?= USD) matches digits before " USD" without including " USD" in the match. Use case: extracting values adjacent to labels without including the label. Only available in PCRE (grep -P), not in sed or awk.	
regex/e5f6a7b8c9d1	regex	easy	regex, patterns, common	What are common regex patterns for matching IPs, emails, and URLs?	IP address: [0-9]+\.[0-9]+\.[0-9]+\.[0-9]+ (simple, allows invalid octets) or more strict with range checks.\nEmail (basic): [a-zA-Z0-9._%+-]+@[a-zA-Z0-9.-]+\.[a-zA-Z]{2,}\nURL: https?://[^\s]+ (simple) or https?://[a-zA-Z0-9.-]+(/[^\s]*)? for more structure.\nThese are pragmatic patterns for log extraction, not RFC-compliant validators.\n\nRemember: "\\b = word boundary." It matches the position between a word character and a non-word character. \\bcat\\b matches "cat" but not "concatenate."	
regex/f6a7b8c9d0e2	regex	medium	regex, grep, pcre-posix	What are the practical differences between PCRE and POSIX regex in everyday use?	POSIX (grep, sed, awk): portable across all Unix, two flavors (BRE and ERE). No lookahead, no lazy quantifiers, no \d shorthand. Use [[:digit:]] instead of \d.\nPCRE (grep -P, Perl, Python): richer features — lookahead/lookbehind, lazy quantifiers, non-capturing groups, \d/\w/\s shortcuts, named groups (?P<name>...).\nRule of thumb: use POSIX for shell scripts that must be portable, PCRE for complex extraction tasks.	
regex/a7b8c9d0e1f3	regex	hard	regex, sed, awk, grep	How do you use regex differently in sed, awk, and grep?	grep: filter lines matching a pattern (grep -E 'ERROR|WARN' log.txt).\nsed: find and replace within lines (sed -E 's/old/new/g' file).\nawk: split fields and apply logic (awk '/ERROR/ {print $1, $NF}' log.txt).\nKey difference: grep selects lines, sed transforms text, awk processes structured data. Combine them: grep filters, sed cleans, awk aggregates. All three support ERE with -E (grep/sed) or natively (awk).	
regex/7504882417a0	regex	medium	regex;lookahead	What is a lookahead assertion in regex and when would you use one?	A lookahead (?=...) matches a position where the pattern ahead matches, without consuming characters. Used to validate constraints (e.g., password must contain a digit) without advancing the cursor.\n\nGotcha: Lookaheads (?=...) and lookbehinds (?<=...) match a position, not characters. They don't consume input.\n\nRemember: "Lookaround = peek without eating."	
regex/281bbff154d3	regex	hard	regex;performance	Why can certain regex patterns cause catastrophic backtracking, and how do you prevent it?	Nested quantifiers like (a+)+ on non-matching input cause exponential backtracking. Prevent by using atomic groups, possessive quantifiers, or restructuring to avoid ambiguous repetition.\n\nRemember: "Backslash-d = digit, backslash-w = word char, backslash-s = space." Uppercase versions negate: \\D = non-digit.	
regex/392358d09a60	regex	medium	regex;groups	Explain the difference between capturing groups and non-capturing groups in regex.	Capturing groups (...) store the matched text for back-references or extraction. Non-capturing groups (?:...) group without storing, which is faster when you only need grouping for alternation or quantifiers.\n\nRemember: "Anchors: ^ = start of line, $ = end of line." They match positions, not characters.	

<!-- wiki:related:start -->
---

## Wiki Navigation

### Related Content

- [Regex & Text Wrangling](../../../../library/topics/regex-text-wrangling/index.md) (Topic Pack, L1) — Regex & Text Wrangling

<!-- wiki:related:end -->
