How to Find Duplicates in a Java Stream (Count and Remove)

Find duplicates in a Java stream with groupingBy() and counting(), list repeated elements, count each one, and remove duplicates from lists or by a field.

One list turned into a frequency map, a list of duplicates, a list of single elements and a list without duplicates

To find duplicates in a Java stream, we count each element with Collectors.groupingBy() and Collectors.counting() and keep the elements whose count is greater than one; to remove duplicates, we call distinct(). The same frequency map also tells us how many times each duplicate appears.

We need these checks when we validate imported files, report repeated orders or user names, or clean a list before saving it. The following example finds, counts and removes duplicate city names.

List<String> cities = List.of("Pune", "Delhi", "Pune", "Goa", "Delhi", "Pune");
Map<String, Long> counts = cities.stream().collect(Collectors.groupingBy(Function.identity(), LinkedHashMap::new, Collectors.counting()));   // {Pune=3, Delhi=2, Goa=1}
List<String> duplicates = counts.entrySet().stream().filter(e -> e.getValue() > 1).map(Map.Entry::getKey).toList();   // [Pune, Delhi]
List<String> withoutDuplicates = cities.stream().distinct().toList();   // [Pune, Delhi, Goa]

Notice the LinkedHashMap::new argument, which keeps the counts in the order of first appearance. We look at the frequency map in detail, two ways to list the duplicates, removing them, duplicates in objects by a field, and which approach to pick.

1. Counting How Many Times Each Element Appears

Almost every duplicate question starts with a frequency map, a Map from each distinct element to the number of times it occurs. Once we have that map, finding duplicates, finding unique elements and reporting counts each take one more short step.

One list turned into a frequency map, a list of duplicates, a list of single elements and a list without duplicates
The frequency map answers every duplicate question, and distinct() removes duplicates without counting

The Collectors.groupingBy() collector takes three arguments here. The classifier Function.identity() uses the element itself as the key, the map factory picks the map type, and the downstream collector counting() counts the elements per key as a Long.

List<Integer> numbers = List.of(1, 1, 2, 3, 3, 3, 4, 5, 6, 6, 6, 7, 8);
Map<Integer, Long> inOrder = numbers.stream().collect(Collectors.groupingBy(Function.identity(), LinkedHashMap::new, Collectors.counting()));   // {1=2, 2=1, 3=3, 4=1, 5=1, 6=3, 7=1, 8=1}
Map<Integer, Long> sortedKeys = numbers.stream().collect(Collectors.groupingBy(Function.identity(), TreeMap::new, Collectors.counting()));      // {1=2, 2=1, 3=3, 4=1, 5=1, 6=3, 7=1, 8=1}
Map<Integer, Long> anyOrder = numbers.stream().collect(Collectors.groupingBy(Function.identity(), Collectors.counting()));
long timesSix = anyOrder.get(6);   // 3

Without a map factory, groupingBy() returns a HashMap, and the printed order of a HashMap is not guaranteed, so we read values by key. A TreeMap sorts the keys, and a LinkedHashMap keeps the source order.

The older form with Collectors.toMap() gives the same counts. It maps every element to 1L and adds the values when a key repeats.

List<String> votes = List.of("red", "blue", "red");
Map<String, Long> tally = votes.stream().collect(Collectors.toMap(Function.identity(), v -> 1L, Long::sum, LinkedHashMap::new));   // {red=2, blue=1}

2. Find Duplicates in a Java Stream

There are two common ways to list the duplicates themselves. The frequency map works for every case and also gives the counts, while a Set lookup finds them in a single pass when we do not need the counts.

2.1. Filtering the Frequency Map

We stream the entries of the frequency map and keep the keys with a count above one. Changing the condition to == 1 gives the elements that appear only once, which is the other half of many reports.

List<String> logins = List.of("ana", "raj", "ana", "lee", "raj", "ana");
Map<String, Long> perUser = logins.stream().collect(Collectors.groupingBy(Function.identity(), LinkedHashMap::new, Collectors.counting()));
List<String> repeated = perUser.entrySet().stream().filter(e -> e.getValue() > 1).map(Map.Entry::getKey).toList();   // [ana, raj]
List<String> once = perUser.entrySet().stream().filter(e -> e.getValue() == 1).map(Map.Entry::getKey).toList();      // [lee]
Map<String, Long> repeatCounts = perUser.entrySet().stream().filter(e -> e.getValue() > 1).collect(Collectors.toMap(Map.Entry::getKey, Map.Entry::getValue, Long::sum, LinkedHashMap::new));   // {ana=3, raj=2}

2.2. Using Set.add() in One Pass

The method Set.add() returns false when the element is already in the set. A filter that keeps the elements for which add() returns false passes every repeated occurrence, and a final distinct() reports each duplicate once.

List<String> logins = List.of("ana", "raj", "ana", "lee", "raj", "ana");
Set<String> seen = new HashSet<>();
List<String> extraOccurrences = logins.stream().filter(name -> !seen.add(name)).toList();   // [ana, raj, ana]
Set<String> seen2 = new HashSet<>();
List<String> duplicateNames = logins.stream().filter(name -> !seen2.add(name)).distinct().toList();   // [ana, raj]

The Set.add() trick changes shared state inside a lambda, so it is safe only in a sequential stream with a new set for every run. In a parallel stream, a plain HashSet can lose updates. The frequency map from section 1 has neither problem, so it is the safer default in shared code.

3. Removing Duplicates From a Stream

The distinct() method removes duplicates and keeps the first occurrence of each element in encounter order. Collecting into a Set also removes duplicates, but a HashSet does not keep the order, whereas a LinkedHashSet does.

List<Integer> withRepeats = List.of(4, 1, 4, 2, 1);
List<Integer> distinctList = withRepeats.stream().distinct().toList();                                   // [4, 1, 2]
Set<Integer> orderedSet = withRepeats.stream().collect(Collectors.toCollection(LinkedHashSet::new));      // [4, 1, 2]
Set<Integer> plainSet = withRepeats.stream().collect(Collectors.toSet());
int setSize = plainSet.size();                                                                           // 3

Both distinct() and sets compare elements with equals() and hashCode(). For our own classes, a record provides both methods, while a plain class that overrides neither keeps every object, because each instance is only equal to itself. To change the original ArrayList in place, we use a LinkedHashSet copy or removeIf() instead of a stream.

4. Duplicates in a List of Objects by a Field

With objects, “duplicate” most often means that one field repeats, not that the whole object is equal. A nightly job imports invoices from a CSV file, and the same invoice number sometimes appears twice because a partner re-sends a row with a corrected amount. The job must report the repeated numbers and import one invoice per number.

record Invoice(String number, String customer, double amount) {}

Grouping by Invoice::number finds the repeated numbers. Collecting with toMap() and a merge function removes the duplicates and lets us choose which row wins. The merge function (first, _) -> first keeps the first row and uses the Java 22 unnamed variable for the second, while (_, last) -> last keeps the corrected row that arrived later.

List<Invoice> rows = List.of(new Invoice("A1", "acme", 100.0), new Invoice("A2", "globex", 75.0), new Invoice("A1", "acme", 110.0), new Invoice("A3", "initech", 20.0));
Map<String, Long> perNumber = rows.stream().collect(Collectors.groupingBy(Invoice::number, LinkedHashMap::new, Collectors.counting()));   // {A1=2, A2=1, A3=1}
List<String> repeatedNumbers = perNumber.entrySet().stream().filter(e -> e.getValue() > 1).map(Map.Entry::getKey).toList();   // [A1]
Map<String, Invoice> latest = rows.stream().collect(Collectors.toMap(Invoice::number, Function.identity(), (_, last) -> last, LinkedHashMap::new));
double a1Amount = latest.get("A1").amount();   // 110.0
int importedCount = latest.size();             // 3

The LinkedHashMap keeps the invoices in file order, so the import log lists the rows in the same order as the CSV. When the duplicate rule uses several fields, such as number and customer, the key becomes a record or a List of those fields, as the article on distinct by multiple fields shows.

5. Which Duplicate Technique to Use

All the stream approaches run in linear time, so the difference is in what they return and whether they are safe in parallel. A per-element Collections.frequency() call, which some tutorials use, scans the whole list for every element and grows with the square of the list size.

GoalCodeNotes
Count every elementgroupingBy(identity(), LinkedHashMap::new, counting())Order of first appearance; safe in parallel
List duplicatesFilter the counts with > 1Works for any type and any key
List duplicates, one passfilter(e -> !seen.add(e))Sequential only; a new set per run
Remove duplicatesdistinct() or LinkedHashSetKeeps the first occurrence
Remove duplicates by fieldtoMap(key, identity(), merge, LinkedHashMap::new)Merge function picks first, last or best
AvoidCollections.frequency(list, e) per elementQuadratic time on large lists

6. Stream Duplicates FAQs

Search results for this topic list a few related tasks that use the same counting technique.

6.1. How Do We Find Duplicate Characters in a String With Streams?

We stream the characters with chars(), box them with mapToObj() and count them the same way. The article on finding duplicate characters covers more variants.

String word = "programming";
List<Character> repeatedChars = word.chars().mapToObj(c -> (char) c).collect(Collectors.groupingBy(Function.identity(), LinkedHashMap::new, Collectors.counting())).entrySet().stream().filter(e -> e.getValue() > 1).map(Map.Entry::getKey).toList();   // [r, g, m]

6.2. How Do We Find the Elements Two Lists Have in Common?

We put one list into a HashSet and filter the other list with contains(), which is a hash lookup. A distinct() at the end removes repeats inside the second list.

List<String> monday = List.of("ana", "raj", "lee");
List<String> tuesday = List.of("lee", "kim", "ana", "lee");
Set<String> mondaySet = new HashSet<>(monday);
List<String> bothDays = tuesday.stream().filter(mondaySet::contains).distinct().toList();   // [lee, ana]

6.3. How Do We Find Duplicates Ignoring Case?

We normalize the key in the classifier, for example with s -> s.toLowerCase(Locale.ROOT) in place of Function.identity(). The counts merge “Pune” and “pune” under one key.

6.4. Can a Stream Count Duplicates in a Parallel Stream?

Yes. The groupingBy() and toMap() collectors build partial maps per thread and merge them, so the counts are correct. For large parallel streams, Collectors.groupingByConcurrent() merges into one ConcurrentMap, which avoids the merge step but gives up the key order.

7. Conclusion

A frequency map built with groupingBy() and counting() answers most duplicate questions in a stream. Filtering it with > 1 lists the duplicates, == 1 lists the unique elements, and a LinkedHashMap or TreeMap factory gives a predictable order.

To remove duplicates, distinct() keeps the first occurrence, and toMap() with a merge function removes duplicates by a field while choosing the surviving element. The Java Stream tutorials list the other collectors.

8. References

Happy Learning !!

Source Code on Github

About Us

HowToDoInJava provides tutorials and how-to guides on Java and related technologies.

It also shares the best practices, algorithms & solutions and frequently asked interview questions.