Regular expressions
Pattern matching for logs, scraped text, and model output — search, findall, groups, and substitution.
Why this matters in AI / ML / GenAI
Extracting JSON from a chatty LLM response, stripping markdown fences, pulling error codes out of logs, redacting emails and phone numbers before they reach a model, and cleaning scraped documents are all regex jobs.
1. Read
Understand the idea in plain English first.
2. Run
Load any example into the compiler and press Run.
3. Change
Edit one value, rerun, and learn from the output.
The core functions
import re, then:
re.search(pattern, text)— first match anywhere, orNonere.match(pattern, text)— must match at the startre.fullmatch— must match the whole stringre.findall— every match as a listre.finditer— every match as objects, with positionsre.sub(pattern, repl, text)— replacere.split(pattern, text)— split on a patternre.compile(pattern)— compile once, reuse in a loop
Always write patterns as raw strings: r"\d+". Without the r, Python interprets the backslash first and the pattern breaks.
A match object is truthy, so if re.search(...) reads naturally. Get the text with .group(), and the position with .start().
Pattern syntax
| Pattern | Matches |
|---|---|
. | any character except newline |
\d \D | digit / non-digit |
\w \W | word character / non-word |
\s \S | whitespace / non-whitespace |
^ $ | start / end of string |
\b | word boundary |
* + ? | 0+, 1+, 0 or 1 |
{2,5} | between 2 and 5 |
[abc] [^abc] | character set / negated |
(...) | capture group |
(?:...) | group without capturing |
(?P<name>...) | named group |
| `a | b` |
Quantifiers are greedy by default — .* grabs as much as possible. Add ? to make them lazy: .*? stops at the first opportunity. Extracting content between two markers almost always needs the lazy form.
Useful flags: re.IGNORECASE, re.MULTILINE (^/$ match each line), and re.DOTALL (. also matches newlines).
When not to use regex
Regex is the wrong tool for nested structures. Do not parse HTML, JSON, or code with it — use json.loads, an HTML parser, or ast. A regex that half-works on nested data fails on the input that matters.
Watch out for catastrophic backtracking: patterns with nested quantifiers such as (a+)+b can take exponential time on adversarial input. Keep patterns simple and anchor them.
Compile patterns used inside loops. re.compile once outside the loop avoids re-parsing the pattern on every iteration.
Hands-on practice
Compiler on the left. Examples on the right.
On desktop, keep the compiler beside the examples. On mobile, the same blocks stack cleanly. Pick an example, try it in the compiler, then change one small thing.
Extract hyperparameters from a training log
Live Python compiler
Try it — in-browser Python
Add a new key=value pair to the log line and confirm it is captured.
Code editor
Output
Python runs in your browser. First run downloads the runtime.
Press Run (or Ctrl+Enter) to execute.
Runs CPython in your browser. NumPy, pandas, scikit-learn and Matplotlib load on demand. Charts appear below the output. No input(), no GPU, no network installs.
Clear code examples
Every example is copy-ready. Use Try in compiler when you want to experiment without scrolling around.
Example
Extract JSON from a chatty LLM reply
The single most common regex task in GenAI code.
import json
import re
reply = '''Sure! Here is the result:
```json
{"sentiment": "positive", "confidence": 0.92}
```
Let me know if you need anything else.'''
fenced = re.search(r"```(?:json)?\s*(\{.*?\})\s*```", reply, re.DOTALL)
if fenced:
data = json.loads(fenced.group(1))
print("from fence:", data)
bare = re.search(r"\{.*\}", reply, re.DOTALL)
print("fallback match:", bare.group()[:40] if bare else None)
print("confidence:", data["confidence"])Example
Parse a log line with named groups
Named groups turn a match into a labelled dict.
import re
pattern = re.compile(
r"(?P<date>\d{4}-\d{2}-\d{2})\s+"
r"(?P<level>DEBUG|INFO|WARNING|ERROR)\s+"
r"(?P<message>.+?)\s+"
r"latency_ms=(?P<latency>\d+)"
)
lines = [
"2026-09-04 ERROR model call failed latency_ms=2450",
"2026-09-04 INFO retrieved 4 chunks latency_ms=87",
"malformed line without fields",
]
for line in lines:
match = pattern.search(line)
if not match:
print("no match:", line)
continue
fields = match.groupdict()
fields["latency"] = int(fields["latency"])
print(fields)Example
Redact personal data before it reaches a model
Substitution with a replacement string, applied in sequence.
import re
text = ("Contact Priya at priya.sharma@example.com or +91 91000 28801. "
"Card 4111-1111-1111-1111, account 998877.")
rules = [
(r"[\w.+-]+@[\w-]+\.[\w.]+", "[EMAIL]"),
(r"\+?\d[\d\s-]{8,}\d", "[PHONE]"),
(r"\b(?:\d{4}[- ]?){3}\d{4}\b", "[CARD]"),
]
redacted = text
for pattern, replacement in rules:
redacted = re.sub(pattern, replacement, redacted)
print("before:", text)
print("after :", redacted)Example
Greedy vs lazy
Greedy .* swallows everything up to the last marker.
import re
text = "<title>First</title> and <title>Second</title>"
greedy = re.findall(r"<title>(.*)</title>", text)
lazy = re.findall(r"<title>(.*?)</title>", text)
print("greedy:", greedy)
print("lazy :", lazy)
print("\nfindall with groups:", re.findall(r"(\w+)=(\d+)", "lr=5 epochs=30 batch=16"))
print("split on punctuation:", re.split(r"[.,;!?]\s*", "One. Two, three; four!"))
print("word boundary:", re.findall(r"\bml\b", "ml mlops ml-ops ml"))Example
Clean scraped text for embedding
A short pipeline of substitutions before chunking.
import re
raw = """ # Heading with **markdown**
Visit https://example.com/docs for more info.
Contact: support@example.com <br/>
Footnote [1] and [2]. """
steps = [
(r"https?://\S+", " "),
(r"[\w.+-]+@[\w-]+\.[\w.]+", " "),
(r"<[^>]+>", " "),
(r"\[\d+\]", " "),
(r"[#*_`]+", " "),
(r"\s+", " "),
]
clean = raw
for pattern, replacement in steps:
clean = re.sub(pattern, replacement, clean)
clean = clean.strip()
print("raw chars :", len(raw))
print("clean chars:", len(clean))
print("clean text :", clean)Takeaways
- Always write patterns as raw strings; compile patterns used in loops.
- Quantifiers are greedy — use .*? when extracting between markers.
- Named groups give you a labelled dict; never parse HTML or JSON with regex.