The regex for alphanumeric characters is ^[a-zA-Z0-9]+$: it accepts a string made only of the letters A to Z (either case) and the digits 0 to 9. The square brackets list the allowed characters, + requires at least one of them, and ^ and $ make the rule apply to the whole string. Spaces, hyphens, the “_” character and accented letters are rejected.
This tutorial explains how to validate alphanumeric input in Java, how to add a length limit, how to require at least one letter and one digit, and how to accept letters from any language. It also shows the common variants (with “_”, a space or a hyphen), how to remove every non-alphanumeric character, and a check that needs no regex.
| To allow | Regex | Java string literal |
|---|---|---|
| Letters and digits only | ^[a-zA-Z0-9]+$ | “^[a-zA-Z0-9]+$” |
| Letters and digits, 3 to 16 characters | ^[a-zA-Z0-9]{3,16}$ | “^[a-zA-Z0-9]{3,16}$” |
| At least one letter and one digit | ^(?=.*[a-zA-Z])(?=.*\d)[a-zA-Z0-9]+$ | “^(?=.*[a-zA-Z])(?=.*\\d)[a-zA-Z0-9]+$” |
| Letters and digits plus “_” | ^\w+$ | “^\\w+$” |
| Letters and digits in any language | (?U)^\p{Alnum}+$ | “(?U)^\\p{Alnum}+$” |
Run against sample strings, these patterns return:
"abc123".matches("[a-zA-Z0-9]+"); // true
"abc 123".matches("[a-zA-Z0-9]+"); // false, space
"".matches("[a-zA-Z0-9]+"); // false, empty
"Lokesh".matches("[a-zA-Z0-9]{3,16}"); // true
"abc123".matches("(?=.*[a-zA-Z])(?=.*\\d)[a-zA-Z0-9]+"); // true
"abc".matches("(?=.*[a-zA-Z])(?=.*\\d)[a-zA-Z0-9]+"); // false, no digit
"caf\u00e9".matches("[a-zA-Z0-9]+"); // false, accented e
"caf\u00e9".matches("(?U)\\p{Alnum}+"); // true
"Hello, World! 123".replaceAll("[^a-zA-Z0-9]", ""); // "HelloWorld123"
1. The Alphanumeric Regex Explained
Alphanumeric characters are letters and digits. In ASCII text, that means 26 lowercase letters, 26 uppercase letters and 10 digits, 62 characters in total. The pattern ^[a-zA-Z0-9]+$ has four parts:
| Part | Meaning |
|---|---|
| ^ | Start of the string |
| [a-zA-Z0-9] | One character: a lowercase letter, an uppercase letter or a digit |
| + | One or more of the previous character class |
| $ | End of the string |
The ranges a-z, A-Z and 0-9 sit inside one character class, so their order does not matter: [0-9A-Za-z] is the same set. The class matches exactly one character; the + after it is what makes the pattern accept a whole word.
2. Validating Alphanumeric Input in Java
In Java, String.matches() is the shortest way to run the check. It returns true only when the regex matches the entire string, so the ^ and $ anchors are optional with it. For code that runs on every request, we compile the pattern once into a static final field:
static final Pattern ALPHANUMERIC = Pattern.compile("^[a-zA-Z0-9]+$");
ALPHANUMERIC.matcher("Lokesh123").matches(); // true
ALPHANUMERIC.matcher("Lokesh123-").matches(); // false
These sample inputs show what the pattern accepts and rejects:
| Input | Result | Reason |
|---|---|---|
| “Lokesh” | true | Letters only |
| “Lokesh123” | true | Letters and digits |
| “2026” | true | Digits only |
| “Lokesh123-“ | false | Hyphen |
| “Lokesh 123” | false | Space |
| “user_1” | false | The “_” character |
| “” | false | Empty string, + needs one character |
| “cafe” with an accented “e” | false | Non-ASCII letter |
Since Java 11, Pattern.asMatchPredicate() turns a pattern into a Predicate<String> that checks the whole string. It fits stream filters and validation APIs that accept a predicate:
Predicate<String> isAlphanumeric = Pattern.compile("[a-zA-Z0-9]+").asMatchPredicate();
isAlphanumeric.test("Lokesh123"); // true
isAlphanumeric.test("Lokesh123-"); // false
For validation, we use matches() or anchors, never a plain find(). find() succeeds when any part of the input matches, so [a-zA-Z0-9]+ finds “Lokesh123” inside “Lokesh123-” and accepts it:
Pattern.compile("[a-zA-Z0-9]+").matcher("Lokesh123-").find(); // true
3. Empty Strings and Length Limits
The quantifier decides whether an empty value passes and how long the value may be:
- * allows zero characters, so [a-zA-Z0-9]* accepts an empty string.
- + requires at least one character.
- {3,16} requires between 3 and 16 characters, a typical rule for user names.
"".matches("[a-zA-Z0-9]*"); // true
"".matches("[a-zA-Z0-9]+"); // false
"ab".matches("[a-zA-Z0-9]{3,16}"); // false, too short
"Lokesh".matches("[a-zA-Z0-9]{3,16}"); // true
"a".repeat(17).matches("[a-zA-Z0-9]{3,16}"); // false, too long
A length limit in the regex keeps the rule in one place. When the error message must say why the value failed, separate checks with String.length() give clearer messages. The article Regex to validate min and max length of input covers length rules in detail.
4. Allowing Spaces, Hyphens and Other Extra Characters
Real input rules often allow one or two more characters. We add them to the character class:
| To allow | Regex | Note |
|---|---|---|
| Letters, digits and “_” | ^\w+$ or ^[a-zA-Z0-9_]+$ | \w is the same set as the class |
| Letters, digits, “_” and “-“ | ^[a-zA-Z0-9_-]+$ | A hyphen at the end of the class is literal |
| Letters, digits and spaces | ^[a-zA-Z0-9 ]+$ | A literal space only |
| Lowercase letters and digits | ^[a-z0-9]+$ | Rejects uppercase |
| Letters and digits, case-insensitive | (?i)^[a-z0-9]+$ | Same result as [a-zA-Z0-9] |
"user_1".matches("\\w+"); // true
"user-1".matches("[a-zA-Z0-9_-]+"); // true
"abc 123".matches("[a-zA-Z0-9 ]+"); // true
"abc\t123".matches("[a-zA-Z0-9 ]+"); // false, tab
"abc\t123".matches("[a-zA-Z0-9\\s]+"); // true, \s allows tabs and line breaks
"Abc123".matches("[a-z0-9]+"); // false, uppercase A
"ABC123".matches("(?i)[a-z0-9]+"); // true
To allow spaces, we write a literal space in the class, not \s. The class \s also accepts tabs and line breaks, which are rarely wanted in a name or a code.
5. Requiring at Least One Letter and One Digit
Some rules, such as product codes or simple passwords, need both kinds of characters. A lookahead checks a condition without consuming text, so we place one lookahead per condition at the start of the pattern:
| Part | Meaning |
|---|---|
| (?=.*[a-zA-Z]) | Somewhere ahead there is a letter |
| (?=.*\d) | Somewhere ahead there is a digit |
| [a-zA-Z0-9]+ | The whole string contains only letters and digits |
String letterAndDigit = "^(?=.*[a-zA-Z])(?=.*\\d)[a-zA-Z0-9]+$";
"abc123".matches(letterAndDigit); // true
"abc".matches(letterAndDigit); // false, no digit
"123".matches(letterAndDigit); // false, no letter
"abc123!".matches(letterAndDigit); // false, "!" is not allowed
The same technique extends to stricter password rules. The article Regex password validator builds on it.
6. Letters and Digits in Any Language
[a-zA-Z0-9] is ASCII-only, so it rejects names such as “cafe” written with an accented “e”. In Java, \p{Alnum} is also ASCII-only by default, despite its name: it is the POSIX class [a-zA-Z0-9]. The UNICODE_CHARACTER_CLASS flag, written inline as (?U), switches \p{Alnum}, \w and \d to their Unicode definitions.
"\u00e9".matches("\\p{Alnum}"); // false, e with an acute accent
"\u00e9".matches("(?U)\\p{Alnum}"); // true
"\u0663".matches("\\d"); // false, Arabic-Indic digit three
"\u0663".matches("(?U)\\p{Alnum}"); // true
"\u00bd".matches("(?U)\\p{Alnum}"); // false, the fraction one half
"\u00bd".matches("[\\p{L}\\p{N}]"); // true
"\u00e9".matches("[\\p{L}\\p{Nd}]"); // true
There are two common Unicode choices, and they differ in what counts as a number:
- (?U)\p{Alnum} accepts alphabetic characters and decimal digits in any script.
- [\p{L}\p{N}] accepts any letter (\p{L}) and any number (\p{N}), which also includes symbols such as the fraction one half and superscript two. [\p{L}\p{Nd}] limits numbers to decimal digits.
For names and user input in several languages, (?U)^\p{Alnum}+$ is a good default.
7. Removing Non-Alphanumeric Characters
To clean a string instead of rejecting it, we negate the class with a caret and replace every match with an empty string. Replacing runs of other characters with a hyphen produces a URL-friendly value.
"Hello, World! 123".replaceAll("[^a-zA-Z0-9]", ""); // "HelloWorld123"
"a-b_c d".replaceAll("[^a-zA-Z0-9]+", "-"); // "a-b-c-d"
For text that contains control characters or non-printable bytes, the article Clean ASCII text of non-printable characters shows the matching character classes.
8. Checking Without a Regex
Character.isLetterOrDigit() checks one character against the Unicode rules, and String.chars() streams the characters of a string. Together they give a regex-free check:
"Lokesh123".chars().allMatch(Character::isLetterOrDigit); // true
"caf\u00e9".chars().allMatch(Character::isLetterOrDigit); // true, Unicode letters pass
"".chars().allMatch(Character::isLetterOrDigit); // true, empty string passes
This check returns true for an empty string, so we add an isEmpty() test when a value is required. It also accepts non-ASCII letters; when only ASCII is allowed, the regex ^[a-zA-Z0-9]+$ is the simpler rule.
9. Trying the Checks Yourself
The alphanumeric project on GitHub prints every check from this article with its result. Its JUnit 6 tests assert the valid and invalid samples, the variants and the Unicode cases. It needs Java 25 and Maven.
mvn -q compile exec:java
mvn test
10. Conclusion
^[a-zA-Z0-9]+$ is the regex for an ASCII alphanumeric string; in Java we run it with matches() or with a compiled Pattern. We change the quantifier for length rules, add characters to the class for variants such as “_” or a space, and add lookaheads when a letter and a digit are both required. For input in other languages, (?U)^\p{Alnum}+$ accepts letters and digits from any script, because plain \p{Alnum} is ASCII-only in Java.
11. References
The POSIX and Unicode character classes used here are listed in the Pattern JavaDoc.
- Pattern (Java SE 25 API)
- Pattern.asMatchPredicate() (Java SE 25 API)
- Character.isLetterOrDigit() (Java SE 25 API)
- Unicode Regular Expressions, UTS #18
Happy Learning !!
But that regular expression won’t work for languages using a different script such as Cyrillic. How do you make a regular expression recognize alphanumeric character in any Java Locale supported language?
String regex = “\\p{Alnum}+”;
This is a Unicode property that matches any alphanumeric character (letters and digits) recognized in any Java Locale.