Regex Word Boundary in Java: Words That Start or End With

How the regex word boundary \b works in Java, with tested patterns for words that start or end with given letters and the mistakes that make \b fail.

Java regex

A regex word boundary, written \b, is a position where a word character meets a non-word character or the edge of the text. It lets a pattern find “cat” as a separate word while it skips the “cat” inside “category”. Like ^ and $, it is an anchor, so it matches a position and consumes no characters.

This tutorial shows how \b and \B decide where a word starts and ends, and how to use them to find words that start or end with given letters. It then combines word boundaries with line anchors to select log lines that begin or end with a specific word, explains the \G anchor, and lists the mistakes that make a \b pattern silently fail in Java.

Each pattern below shows its result as a comment:

Pattern.compile("\\bcat\\b").matcher("the cat sat").find();     // true
Pattern.compile("\\bcat\\b").matcher("category").find();        // false

// find every match in a text
Pattern.compile("\\bun\\w*").matcher("undo the unknown fun").results()
    .map(MatchResult::group).toList();                          // [undo, unknown]
Pattern.compile("\\w+ing\\b").matcher("walking and talking in the morning").results()
    .map(MatchResult::group).toList();                          // [walking, talking, morning]
Pattern.compile("\\Bcat\\B").matcher("concatenate cat").results()
    .map(MatchResult::group).toList();                          // [cat], the one inside "concatenate"
To findRegexJava string literal
The whole word “cat”\bcat\b“\\bcat\\b”
Words starting with “un”\bun\w*“\\bun\\w*”
Words ending with “ing”\w+ing\b“\\w+ing\\b”
“cat” inside a longer word\Bcat\B“\\Bcat\\B”
Lines starting with the word “ERROR”(?m)^ERROR\b.*“(?m)^ERROR\\b.*”

1. The Three Positions Where \b Matches

A word character is a letter, a digit or the “_” character, the same set as \w. The anchor \b matches at three kinds of positions:

  1. Before the first character of the text, if that character is a word character.
  2. After the last character of the text, if that character is a word character.
  3. Between two characters where one is a word character and the other is not.

Replacing each boundary with a “|” makes the positions visible. In “cat sat”, the boundaries sit on both sides of each word, and the space has a boundary on each side:

"cat sat".replaceAll("\\b", "|");        // "|cat| |sat|"
"a cat".replaceAll("\\B", "|");          // "a c|a|t"

The anchor \B is the opposite of \b: it matches at every position that is not a word boundary. In “a cat”, those are the positions between two letters of “cat”.

A boundary before a word character is the start of a word, and a boundary after a word character is the end of a word. So \b\w finds the first letter of each word and \w\b finds the last one:

Pattern.compile("\\b\\w").matcher("hello big world").results()
    .map(MatchResult::group).toList();      // [h, b, w]
Pattern.compile("\\w\\b").matcher("hello big world").results()
    .map(MatchResult::group).toList();      // [o, g, d]

2. Hyphens, Apostrophes and Accented Letters

\b follows the regex definition of a word, which is not the English one. Only [a-zA-Z0-9_] count as word characters by default, so a hyphen or an apostrophe splits a word into two:

"e-mail".replaceAll("\\b", "|");         // "|e|-|mail|"
"don't".replaceAll("\\b", "|");          // "|don|'|t|"
"user_1".replaceAll("\\b", "|");         // "|user_1|", "_" and digits are word characters

As a result, \bmail\b finds “mail” inside “e-mail”, and \bt\b finds the “t” of “don’t”. When a hyphen should count as part of a word, we replace \b with lookaround that treats the hyphen as a word character:

Pattern.compile("\\bmail\\b").matcher("my e-mail").find();                // true
Pattern.compile("(?<![\\w-])mail(?![\\w-])").matcher("my e-mail").find();  // false
Pattern.compile("(?<![\\w-])mail(?![\\w-])").matcher("mail box").find();   // true

The lookbehind (?<![\w-]) means “not preceded by a word character or a hyphen”, and the lookahead (?![\w-]) means the same for the next character. Section 7 uses the same technique for words that end with a symbol.

Accented letters are not word characters by default either. Since Java 19, \b uses the same ASCII definition as \w (JDK-8264160). In the French word “ete” written with two accented “e” letters (U+00E9), only the plain “t” counts as a word. The flag UNICODE_CHARACTER_CLASS, inline (?U), makes both \w and \b follow Unicode rules:

String ete = "\u00e9t\u00e9";

Pattern.compile("\\b\\w+\\b").matcher(ete).results()
    .map(MatchResult::group).toList();      // [t]
Pattern.compile("(?U)\\b\\w+\\b").matcher(ete).results()
    .map(MatchResult::group).toList();      // [ete] with the accents, the whole word

For text in languages other than English, we add (?U) to every pattern that uses \b or \w.

3. Words That Start With Given Letters

To find words that start with a prefix, we put \b before the prefix and \w* after it. The \b makes the prefix the start of a word, and \w* takes the rest of the word.

Pattern.compile("\\bun\\w*").matcher("undo the unknown fun").results()
    .map(MatchResult::group).toList();      // [undo, unknown]

Pattern.compile("un\\w*").matcher("undo the unknown fun").results()
    .map(MatchResult::group).toList();      // [undo, unknown, un], "un" from "fun"

Pattern.compile("(?i)\\bun\\w*").matcher("Undo the unknown fun").results()
    .map(MatchResult::group).toList();      // [Undo, unknown]

The second pattern has no \b, so it also matches the “un” at the end of “fun”. The third pattern adds (?i) to ignore case. A character class works as the prefix too: \b[A-Z]\w* finds the words that start with a capital letter.

Pattern.compile("\\b[A-Z]\\w*").matcher("Lokesh lives in Delhi").results()
    .map(MatchResult::group).toList();      // [Lokesh, Delhi]

4. Words That End With Given Letters

For a suffix, the boundary goes after it: \w+ing\b. We use \w+ instead of \w* when the suffix alone should not count as a word, so “ing” by itself is not a match.

Pattern.compile("\\w+ing\\b").matcher("walking and talking in the morning").results()
    .map(MatchResult::group).toList();      // [walking, talking, morning]

Pattern.compile("\\w+ing\\b").matcher("singing birds and a kingdom").results()
    .map(MatchResult::group).toList();      // [singing]

Pattern.compile("\\w+ing").matcher("singing birds and a kingdom").results()
    .map(MatchResult::group).toList();      // [singing, king], "king" from "kingdom"

Without the final \b, the pattern returns “king” from “kingdom”, a word that does not end with “ing”.

5. Lines That Start or End With a Word

A common task is to filter log lines by the word they start with. The line anchors ^ and $ with MULTILINE mode select the line, and \b makes sure the first word is exactly the one we want:

String log = """
    ERROR disk full
    INFO started
    ERRORS: 0
    WARN retry done""";

Pattern.compile("(?m)^ERROR.*").matcher(log).results()
    .map(MatchResult::group).toList();      // [ERROR disk full, ERRORS: 0]

Pattern.compile("(?m)^ERROR\\b.*").matcher(log).results()
    .map(MatchResult::group).toList();      // [ERROR disk full]

Pattern.compile("(?m)^.*\\bdone$").matcher(log).results()
    .map(MatchResult::group).toList();      // [WARN retry done]

The first pattern also returns “ERRORS: 0”, because “ERRORS” starts with “ERROR”. The \b after “ERROR” requires a non-word character next, so “ERRORS” no longer matches. In the last pattern, \b before “done” rejects a line that ends with, for example, “undone”. The article Regex start and end of line explains ^, $ and MULTILINE in detail.

6. Matching Inside a Word With \B

\B requires the pattern to be part of a longer word. \Bcat\B finds “cat” only when there are word characters on both sides, and \Bcat finds “cat” at the end or in the middle of a word, but not at its start.

Pattern.compile("\\Bcat\\B").matcher("concatenate cat").results()
    .map(MatchResult::group).toList();      // [cat], from "concatenate"

Pattern.compile("\\Bcat").matcher("concat").find();     // true
Pattern.compile("\\Bcat").matcher("cat").find();        // false

7. Mistakes That Make \b Fail

Most “my word boundary does not work” problems in Java come from one of these three causes.

A single backslash in the Java string. In a Java string literal, “\b” is the backspace character (U+0008), not a regex word boundary. The code compiles, the pattern searches for a backspace, and it never matches normal text. The regex needs “\\b”.

"\b".equals("\u0008");                                       // true
Pattern.compile("\bcat\b").matcher("the cat").find();        // false, searches for backspaces
Pattern.compile("\\bcat\\b").matcher("the cat").find();      // true

A word that starts or ends with a symbol. \b needs a word character on one side. In “I like C++ code”, the “+” and the space after it are both non-word characters, so there is no boundary after “C++” and \bC\+\+\b fails. Lookaround asserts “no word character before or after” and works for any first and last character:

Pattern.compile("\\bC\\+\\+\\b").matcher("I like C++ code").find();            // false
Pattern.compile("(?<!\\w)C\\+\\+(?!\\w)").matcher("I like C++ code").find();   // true
Pattern.compile("(?<!\\w)C\\+\\+(?!\\w)").matcher("I like C++11").find();      // false

Using matches() to search. String.matches() tests the whole input, so “the cat”.matches(“\\bcat\\b”) returns false. To search inside a text, we use Matcher.find(). The article Regex to match an exact word builds whole-word search on top of this, including words taken from user input.

8. The \G Anchor: Continuing From the Last Match

\G is another position anchor. It matches where the previous match ended, or at the start of the input for the first search. Each new match must start right where the last one stopped, so the search stops at the first gap.

Pattern.compile("\\G\\d").matcher("123abc456").results()
    .map(MatchResult::group).toList();      // [1, 2, 3]
Pattern.compile("\\d").matcher("123abc456").results()
    .map(MatchResult::group).toList();      // [1, 2, 3, 4, 5, 6]
Pattern.compile("\\Gdog").matcher("dog dog").results()
    .map(MatchResult::group).toList();      // [dog], the space breaks the chain

\G is useful in tokenizers that read a string piece by piece and must reject anything between the pieces.

9. Word Boundary Code and Tests

The word boundary project on GitHub runs every pattern from this article and prints the result, and its JUnit 6 tests assert each one. It uses Java 25 and Maven.

mvn -q compile exec:java
mvn test

10. Conclusion

\b marks the edge between a word character and anything else, which makes it the tool for whole words, prefixes (\bun\w*) and suffixes (\w+ing\b). Combined with ^, $ and MULTILINE, it selects lines that begin or end with an exact word. In Java code, we write it as “\\b”, add (?U) for non-English text, and switch to lookaround when the word starts or ends with a symbol. The Java regex metacharacters reference lists the other anchors and classes used here.

11. References

The boundary matchers are documented in the Pattern JavaDoc; the other links explain the rules in more depth.

Happy Learning !!

Source Code on Github

Leave a Comment

  1. I have a doubt. Imagine, In a file, if i want all the data that start in a String and ends whit another string, what can i do ?
    Thank you

Comments are closed.

About Us

HowToDoInJava provides tutorials and how-to guides on Java and related technologies.

It also shares the best practices, algorithms & solutions and frequently asked interview questions.