Matching Unicode with \p{...} Property Escapes
\p{L} matches any Unicode letter — far better than [A-Za-z] for real-world text. JavaScript needs the u flag; Python's built-in re does not support \p at all. Details below.
Support at a glance
| Engine | \p{...}? | How to enable |
|---|---|---|
| JavaScript | Yes | Requires the u (or v) flag: /\p{L}/u |
Python re | No | Not supported — use the PyPI regex module |
Python regex | Yes | regex.findall(r'\p{L}+', text) |
| Java | Yes | \p{L}, \p{IsGreek} work by default |
| PHP / PCRE | Yes | Add the u modifier: /\p{L}/u |
| Go (RE2) | Yes | \p{L}, \p{Greek} supported |
The most useful classes
| Escape | Matches |
|---|---|
\p{L} | Any letter (any language, any case) |
\p{N} | Any numeric character |
\p{Lu} / \p{Ll} | Uppercase / lowercase letter |
\p{Script=Greek} | Characters in the Greek script |
\P{L} | Anything that is NOT a letter (capital P negates) |
Use it in your code
// JavaScript — the u flag is mandatory for \p{...}
'Café déjà ü'.match(/\p{L}+/gu); // -> ['Café', 'déjà', 'ü']
# Python — built-in re has no \p; use the regex module
import regex
regex.findall(r'\p{L}+', 'Café déjà ü') # -> ['Café', 'déjà', 'ü']
Why not just [A-Za-z]?
[A-Za-z] silently drops every accented, Cyrillic, Greek, Arabic or CJK character — a bug that only shows up once real users type their names. \p{L} matches letters in every script, so validation and tokenization keep working for international input. Reach for property escapes any time your text is not guaranteed ASCII.
FAQ
Why doesn't \p{L} work in my JavaScript regex?
Unicode property escapes require the u flag (or v flag). Write /\p{L}/u — without it, \p is treated as a literal 'p' in most engines.
Does Python's re module support \p{L}?
No. The built-in re module has no Unicode property escapes. Install the PyPI 'regex' module, which supports \p{L} and \p{Script=...}.
How do I match a specific script like Greek or Cyrillic?
Use \p{Script=Greek} (or \p{Greek} in some engines) and \p{Script=Cyrillic}. Java uses \p{IsGreek}.