Regex to Match the Trademark Symbol in Java (U+2122)

Match the trademark sign U+2122 in Java with \u2122, \x{2122} or \N{TRADE MARK SIGN}, capture brand names, handle the registered and copyright signs, and convert (TM) and HTML entities to the real symbol.

Java regex

The trademark symbol is the Unicode character U+2122 (TRADE MARK SIGN), and in a Java regex we match it with the escape \u2122, written as “\\u2122” in a string literal. Java also accepts \x{2122} and the character name \N{TRADE MARK SIGN}. Writing an escape keeps the source file plain ASCII, so the pattern works whatever encoding the editor, the compiler or the build server uses.

This guide covers finding the trademark sign and its position, capturing the brand name in front of it, matching the related registered, copyright and service mark signs, and converting other spellings such as “(TM)” or the HTML entity into the real symbol. It also points out three mistakes: escape syntax from other languages, a word boundary after the symbol, and the emoji form of the sign. Every result was checked on Java 25.

The sample text contains two trademark signs, and each comment shows the value the code returns.

String text = "Java\u2122 and Duke\u2122 rule";
Pattern tm = Pattern.compile("\\u2122");

"Java\u2122".matches("Java\\u2122");                    // true
tm.matcher(text).find();                                // true
tm.matcher("Java").find();                              // false
tm.matcher(text).results().count();                    // 2

Pattern.compile("(\\w+)\\s?\\u2122").matcher(text).results()
    .map(r -> r.group(1)).toList();                     // [Java, Duke]

text.replaceAll("\\u2122", "");                         // "Java and Duke rule"
"Java(TM)".replaceAll("\\((?i:tm)\\)", "\u2122");       // "Java" followed by the trademark sign

1. Finding the Trademark Symbol and Its Index

The trademark sign is a single char in Java, so a pattern with only this character finds each occurrence. Matcher.results() returns the matches with their start() and end() indexes, where end() is the index after the match.

tm.matcher(text).results()
    .map(r -> r.start() + "-" + r.end())
    .toList();                          // [4-5, 14-15]

For a single fixed character, the String methods are enough and faster than a regex:

text.indexOf('\u2122');                 // 4
text.contains("\u2122");                // true

A regex becomes useful when the symbol is part of a larger rule, such as “a word followed by the trademark sign” or “any of the legal symbols”, as in the next sections.

2. Writing U+2122 in a Java Regex

In Java, the trademark sign can reach the regex engine in several forms, and each of these patterns matches it:

  • the character itself, typed into the source file or created by the Java compiler from “\u2122” (one backslash)
  • the regex escape \u2122, written “\\u2122” in Java (two backslashes); the regex engine reads the escape
  • \x{2122}, the hexadecimal code point form, available since Java 7
  • \N{TRADE MARK SIGN}, the Unicode character name, available since Java 9
  • \p{So}, the category “Symbol, other”, which also matches many other symbols

\p{Sc} does not match it: the trademark sign is not a currency symbol (for those, see the currency symbol regex).

Escape forms from other languages fail in Java, and one of them fails without an error:

"\u2122".matches("\\x2122");      // false, this is \x21 ("!") followed by "22"
"!22".matches("\\x2122");         // true
Pattern.compile("\\u{2122}");      // PatternSyntaxException

The \x form without braces takes exactly two hex digits, so \x2122 means the character “!” (hex 21) followed by the text “22”. The braces form \u{2122} belongs to JavaScript (with the u flag) and Ruby. Java rejects it with this message:

java.util.regex.PatternSyntaxException: Illegal Unicode escape sequence near index 2
\u{2122}
  ^

3. Capturing the Brand Name Before the Symbol

Product names are usually written as a word immediately followed by the trademark sign. The pattern (\w+)\s?\u2122 captures that word in group 1 and allows one optional space before the sign.

Pattern brand = Pattern.compile("(\\w+)\\s?\\u2122");

brand.matcher("Java\u2122 and Duke\u2122 rule").results()
    .map(r -> r.group(1) + " at " + r.start())
    .toList();                                          // [Java at 0, Duke at 10]
brand.matcher("Duke \u2122").results()
    .map(r -> r.group(1)).toList();                     // [Duke]

A word boundary \b after the trademark sign fails at the end of a word. \b marks a position between a word character and a non-word character. The trademark sign is not a word character, so in “Java” followed by the sign at the end of the text, there is no boundary after the sign:

Pattern.compile("\\bJava\\u2122\\b").matcher("Java\u2122").find();   // false
Pattern.compile("\\bJava\\u2122").matcher("Java\u2122").find();      // true

The word boundary article explains this rule in detail. To say “not followed by a letter”, we use the negative lookahead (?!\w) after the symbol.

4. Trademark, Registered and Copyright Symbols Together

Legal texts and product pages mix several related symbols. Their meaning differs: the trademark sign marks a claimed, often unregistered mark, the registered sign marks a mark registered with a trademark office such as the USPTO, and the copyright sign marks creative works.

SymbolCode pointEscape in a Java regexUsed for
Trademark signU+2122\u2122Trademark, registered or not
Registered signU+00AE\u00aeRegistered trademark
Copyright signU+00A9\u00a9Copyright
Service markU+2120\u2120Trademark for a service
Sound recording copyrightU+2117\u2117Copyright of a recording

A character class with all five escapes matches any of them. It is more precise than \p{So}, which also matches arrows, emoji and many other symbols.

Pattern marks = Pattern.compile("[\\u2122\\u00ae\\u00a9\\u2120\\u2117]");
String legal = "Java\u2122, Duke\u00ae, \u00a9 2026";

marks.matcher(legal).results().map(r -> r.group()).toList();   // [trademark, registered, copyright]
marks.matcher(legal).replaceAll("");                           // "Java, Duke,  2026"

The removal leaves two spaces before “2026”. When the cleaned text is shown to users, we also collapse spaces, for example with replaceAll(” {2,}”, ” “).

5. Other Spellings of the Trademark Sign

Plain-text and HTML sources often write the sign without the Unicode character. We convert these spellings to U+2122 first, so later searches need only one pattern:

  • “(TM)” or “(tm)” in plain text, sometimes after a space
  • the HTML named entity ™
  • the numeric entities ™ (decimal) and ™ (hex)
Pattern spellings = Pattern.compile("\\s?\\((?i:tm)\\)|&(?:trade|#8482|#x2122);");

spellings.matcher("Java(TM), Duke (tm), Kotlin™ and Scala™").replaceAll("\u2122");
// each brand now ends with the trademark sign

The inline group (?i:tm) turns on case-insensitive matching only for the letters “tm”. For HTML documents, a parser such as Jsoup decodes all entities at once, which is more reliable than a regex over raw HTML.

Two more forms can appear in text from phones and from Unicode normalization:

  • Emoji style. Text from emoji pickers can contain the variation selector U+FE0F after the sign, which asks for a colored emoji. A find() for \u2122 still succeeds, but a whole-input check needs \u2122\ufe0f?.
  • Compatibility normalization. The NFKC form replaces the trademark sign with the letters “TM” and the service mark with “SM”. The registered and copyright signs stay unchanged. Text that went through Normalizer with NFKC, for example in the ASCII cleaning method, no longer contains U+2122.
"\u2122\ufe0f".matches("\\u2122");                      // false
"\u2122\ufe0f".matches("\\u2122\\ufe0f?");              // true
Pattern.compile("\\u2122").matcher("\u2122\ufe0f").find();   // true

Normalizer.normalize("\u2122", Normalizer.Form.NFKC);   // "TM"
Normalizer.normalize("\u2120", Normalizer.Form.NFKC);   // "SM"
Normalizer.normalize("\u00ae", Normalizer.Form.NFKC);   // registered sign, unchanged

6. Source Code

The trademark symbol project on GitHub contains each pattern from this page and prints its result, with symbols shown as escapes. JUnit 6 tests check all values, including the exception message for \u{2122}. Build it with Java 25 and Maven.

mvn -q compile exec:java
mvn test

7. Conclusion

To match the trademark sign in Java, we use the escape \u2122 (Java string “\\u2122”), or \x{2122} or \N{TRADE MARK SIGN}; braces-only \u{2122} and the short \x2122 do not work. A capturing group in front of the sign extracts the brand name, and a character class covers the registered, copyright and service mark signs as well. Before searching, we convert “(TM)” and HTML entities to the real symbol, and we keep in mind that NFKC normalization turns the sign into the letters “TM”.

8. References

The Pattern JavaDoc documents the \u, \x{…} and \N{…} escapes; the Unicode chart lists the Letterlike Symbols block that contains U+2122.

Happy Learning !!

Source Code on Github

About Us

HowToDoInJava provides tutorials and how-to guides on Java and related technologies.

It also shares the best practices, algorithms & solutions and frequently asked interview questions.