A regular expression (regex) is a compact pattern-matching language for describing sets of strings — used to search, validate, split, or extract text. Nearly every programming language ships a regex engine, and while dialects differ slightly, the core syntax (literals, character classes, quantifiers, anchors) is shared across almost all of them.
A regex is built from a small set of building blocks:
cat matches "cat").[abc] matches one of a, b, or c; \d, \w, \s are shorthand for digits, word characters, and whitespace.* (0+), + (1+), ? (0 or 1), {n,m} (between n and m repetitions) control how many times the preceding token repeats.^ and $ pin a match to the start or end of a line/string; \b marks a word boundary.(...) groups a sub-pattern (see capture group); | means "or".Most engines are backtracking-based: they try a path, and if it fails, they backtrack and try another. This makes some patterns — especially nested quantifiers like (a+)+ — vulnerable to catastrophic backtracking, where a crafted input causes exponential-time matching (a common denial-of-service vector, ReDoS).
Example:
Pattern: ^[\w.+-]+@[\w-]+\.[a-zA-Z]{2,}$
Matches: "user.name+tag@example.com"
. matches almost any character by default — remember to escape it (\.) when you mean a literal dot..) grab as much as possible before backtracking; use the lazy form (.?) when you want the shortest match.