The trademark symbol is the Unicode character U+2122 (TRADE MARK SIGN), and in a Java regex we match it with the escape \u2122, written as “\\u2122” in a string literal. Java also accepts \x{2122} and the character name \N{TRADE MARK SIGN}. Writing an escape keeps the source file plain ASCII, so the pattern works whatever encoding the editor, the compiler or the build server uses.
This guide covers finding the trademark sign and its position, capturing the brand name in front of it, matching the related registered, copyright and service mark signs, and converting other spellings such as “(TM)” or the HTML entity into the real symbol. It also points out three mistakes: escape syntax from other languages, a word boundary after the symbol, and the emoji form of the sign. Every result was checked on Java 25.
The sample text contains two trademark signs, and each comment shows the value the code returns.
String text = "Java\u2122 and Duke\u2122 rule";
Pattern tm = Pattern.compile("\\u2122");
"Java\u2122".matches("Java\\u2122"); // true
tm.matcher(text).find(); // true
tm.matcher("Java").find(); // false
tm.matcher(text).results().count(); // 2
Pattern.compile("(\\w+)\\s?\\u2122").matcher(text).results()
.map(r -> r.group(1)).toList(); // [Java, Duke]
text.replaceAll("\\u2122", ""); // "Java and Duke rule"
"Java(TM)".replaceAll("\\((?i:tm)\\)", "\u2122"); // "Java" followed by the trademark sign
1. Finding the Trademark Symbol and Its Index
The trademark sign is a single char in Java, so a pattern with only this character finds each occurrence. Matcher.results() returns the matches with their start() and end() indexes, where end() is the index after the match.
tm.matcher(text).results()
.map(r -> r.start() + "-" + r.end())
.toList(); // [4-5, 14-15]
For a single fixed character, the String methods are enough and faster than a regex:
text.indexOf('\u2122'); // 4
text.contains("\u2122"); // true
A regex becomes useful when the symbol is part of a larger rule, such as “a word followed by the trademark sign” or “any of the legal symbols”, as in the next sections.
2. Writing U+2122 in a Java Regex
In Java, the trademark sign can reach the regex engine in several forms, and each of these patterns matches it:
- the character itself, typed into the source file or created by the Java compiler from “\u2122” (one backslash)
- the regex escape \u2122, written “\\u2122” in Java (two backslashes); the regex engine reads the escape
- \x{2122}, the hexadecimal code point form, available since Java 7
- \N{TRADE MARK SIGN}, the Unicode character name, available since Java 9
- \p{So}, the category “Symbol, other”, which also matches many other symbols
\p{Sc} does not match it: the trademark sign is not a currency symbol (for those, see the currency symbol regex).
Escape forms from other languages fail in Java, and one of them fails without an error:
"\u2122".matches("\\x2122"); // false, this is \x21 ("!") followed by "22"
"!22".matches("\\x2122"); // true
Pattern.compile("\\u{2122}"); // PatternSyntaxException
The \x form without braces takes exactly two hex digits, so \x2122 means the character “!” (hex 21) followed by the text “22”. The braces form \u{2122} belongs to JavaScript (with the u flag) and Ruby. Java rejects it with this message:
java.util.regex.PatternSyntaxException: Illegal Unicode escape sequence near index 2
\u{2122}
^
3. Capturing the Brand Name Before the Symbol
Product names are usually written as a word immediately followed by the trademark sign. The pattern (\w+)\s?\u2122 captures that word in group 1 and allows one optional space before the sign.
Pattern brand = Pattern.compile("(\\w+)\\s?\\u2122");
brand.matcher("Java\u2122 and Duke\u2122 rule").results()
.map(r -> r.group(1) + " at " + r.start())
.toList(); // [Java at 0, Duke at 10]
brand.matcher("Duke \u2122").results()
.map(r -> r.group(1)).toList(); // [Duke]
A word boundary \b after the trademark sign fails at the end of a word. \b marks a position between a word character and a non-word character. The trademark sign is not a word character, so in “Java” followed by the sign at the end of the text, there is no boundary after the sign:
Pattern.compile("\\bJava\\u2122\\b").matcher("Java\u2122").find(); // false
Pattern.compile("\\bJava\\u2122").matcher("Java\u2122").find(); // true
The word boundary article explains this rule in detail. To say “not followed by a letter”, we use the negative lookahead (?!\w) after the symbol.
4. Trademark, Registered and Copyright Symbols Together
Legal texts and product pages mix several related symbols. Their meaning differs: the trademark sign marks a claimed, often unregistered mark, the registered sign marks a mark registered with a trademark office such as the USPTO, and the copyright sign marks creative works.
| Symbol | Code point | Escape in a Java regex | Used for |
|---|---|---|---|
| Trademark sign | U+2122 | \u2122 | Trademark, registered or not |
| Registered sign | U+00AE | \u00ae | Registered trademark |
| Copyright sign | U+00A9 | \u00a9 | Copyright |
| Service mark | U+2120 | \u2120 | Trademark for a service |
| Sound recording copyright | U+2117 | \u2117 | Copyright of a recording |
A character class with all five escapes matches any of them. It is more precise than \p{So}, which also matches arrows, emoji and many other symbols.
Pattern marks = Pattern.compile("[\\u2122\\u00ae\\u00a9\\u2120\\u2117]");
String legal = "Java\u2122, Duke\u00ae, \u00a9 2026";
marks.matcher(legal).results().map(r -> r.group()).toList(); // [trademark, registered, copyright]
marks.matcher(legal).replaceAll(""); // "Java, Duke, 2026"
The removal leaves two spaces before “2026”. When the cleaned text is shown to users, we also collapse spaces, for example with replaceAll(” {2,}”, ” “).
5. Other Spellings of the Trademark Sign
Plain-text and HTML sources often write the sign without the Unicode character. We convert these spellings to U+2122 first, so later searches need only one pattern:
- “(TM)” or “(tm)” in plain text, sometimes after a space
- the HTML named entity ™
- the numeric entities ™ (decimal) and ™ (hex)
Pattern spellings = Pattern.compile("\\s?\\((?i:tm)\\)|&(?:trade|#8482|#x2122);");
spellings.matcher("Java(TM), Duke (tm), Kotlin™ and Scala™").replaceAll("\u2122");
// each brand now ends with the trademark sign
The inline group (?i:tm) turns on case-insensitive matching only for the letters “tm”. For HTML documents, a parser such as Jsoup decodes all entities at once, which is more reliable than a regex over raw HTML.
Two more forms can appear in text from phones and from Unicode normalization:
- Emoji style. Text from emoji pickers can contain the variation selector U+FE0F after the sign, which asks for a colored emoji. A find() for \u2122 still succeeds, but a whole-input check needs \u2122\ufe0f?.
- Compatibility normalization. The NFKC form replaces the trademark sign with the letters “TM” and the service mark with “SM”. The registered and copyright signs stay unchanged. Text that went through Normalizer with NFKC, for example in the ASCII cleaning method, no longer contains U+2122.
"\u2122\ufe0f".matches("\\u2122"); // false
"\u2122\ufe0f".matches("\\u2122\\ufe0f?"); // true
Pattern.compile("\\u2122").matcher("\u2122\ufe0f").find(); // true
Normalizer.normalize("\u2122", Normalizer.Form.NFKC); // "TM"
Normalizer.normalize("\u2120", Normalizer.Form.NFKC); // "SM"
Normalizer.normalize("\u00ae", Normalizer.Form.NFKC); // registered sign, unchanged
6. Source Code
The trademark symbol project on GitHub contains each pattern from this page and prints its result, with symbols shown as escapes. JUnit 6 tests check all values, including the exception message for \u{2122}. Build it with Java 25 and Maven.
mvn -q compile exec:java
mvn test
7. Conclusion
To match the trademark sign in Java, we use the escape \u2122 (Java string “\\u2122”), or \x{2122} or \N{TRADE MARK SIGN}; braces-only \u{2122} and the short \x2122 do not work. A capturing group in front of the sign extracts the brand name, and a character class covers the registered, copyright and service mark signs as well. Before searching, we convert “(TM)” and HTML entities to the real symbol, and we keep in mind that NFKC normalization turns the sign into the letters “TM”.
8. References
The Pattern JavaDoc documents the \u, \x{…} and \N{…} escapes; the Unicode chart lists the Letterlike Symbols block that contains U+2122.
- Pattern (Java SE 25 API)
- Normalizer (Java SE 25 API)
- Unicode code chart: Letterlike Symbols U+2100 to U+214F (PDF)
- HTML Standard: named character references
- What is a trademark? (USPTO)
Happy Learning !!