To make a Java Stream distinct by multiple fields, we pass a distinctByKey() predicate to filter(). The predicate builds a key from those fields, for example the title and the artist of a song, and remembers every key it has seen, so the first song with a new key passes and later songs with the same key are dropped. A key made of several field values is called a composite key, and we can build it with List.of(title, artist) or a small record. When the last song must win instead, we use Collectors.toMap() with a merge function.
We need distinct by multiple fields when two objects count as duplicates even though some of their fields differ, such as the same song added twice to a playlist with different play counts.
The following example removes duplicate songs by title and artist, with the result of each stream as a comment.
public static <T> Predicate<T> distinctByKey(Function<? super T, ?> keyExtractor) {
Set<Object> seen = ConcurrentHashMap.newKeySet();
return t -> seen.add(keyExtractor.apply(t)); // add() returns false for a key seen before
}
// songs = [rain/anna/120, home/ben/90, rain/anna/75, home/carl/60, home/ben/40, rain/anna/75]
List<Song> byList = songs.stream()
.filter(distinctByKey(s -> List.of(s.title(), s.artist())))
.toList(); // [rain/anna/120, home/ben/90, home/carl/60]
List<Song> byRecord = songs.stream()
.filter(distinctByKey(SongKey::new)) // record SongKey(String title, String artist)
.toList(); // [rain/anna/120, home/ben/90, home/carl/60]
List<Song> keepLast = songs.stream()
.collect(Collectors.collectingAndThen(
Collectors.toMap(SongKey::new, Function.identity(), (a, b) -> b, LinkedHashMap::new),
m -> new ArrayList<>(m.values()))); // [rain/anna/75, home/ben/40, home/carl/60]
Notice that byList and byRecord keep the first song for each key, whereas keepLast keeps the last one.
Next, we look at why plain distinct() is not enough and build the predicate step by step. After that, we cover sorted results with a TreeSet, parallel streams and the Java 24 Gatherer API.
1. Why distinct() Compares All Fields
Stream.distinct() removes duplicates, and it decides whether two elements are equal by calling their equals() and hashCode() methods. A Java record generates both methods for us, and the generated methods compare every field. So two songs are equal only when the title, the artist and the play count all match.
In the examples, we use one list of songs in a playlist, where the same song (same title and artist) appears several times with different play counts.
record Song(String title, String artist, int plays) {
@Override
public String toString() { return title + "/" + artist + "/" + plays; }
}
static List<Song> songs = List.of(
new Song("rain", "anna", 120),
new Song("home", "ben", 90),
new Song("rain", "anna", 75),
new Song("home", "carl", 60),
new Song("home", "ben", 40),
new Song("rain", "anna", 75));
The list is a static field, so the snippets in the following sections can read it from a main() method. Calling distinct() on it removes exact copies only.
List<Song> exactCopiesRemoved = songs.stream().distinct().toList();
// [rain/anna/120, home/ben/90, rain/anna/75, home/carl/60, home/ben/40]
The distinct() call removed only the last rain/anna/75, because that song is the only exact copy of an earlier song. For strings and boxed numbers, distinct() works the same way, because String and the wrapper classes compare their values in equals().
We want to compare only title and artist, and Java gives us two ways to do that.
- Override equals() and hashCode() so they use only those two fields, following the equals() and hashCode() contract. The new rule applies to the whole application, so every HashSet and HashMap that holds songs follows it too.
- Leave the class as it is, and build a key from the two fields inside one stream pipeline. The Song class stays unchanged, so we use the key approach in the examples.
2. Find Elements Distinct by Multiple Fields
Java has no built-in distinctBy() method on Stream, so we write a small helper method, distinctByKey(), that returns a Predicate (a function that returns true or false). The filter() method keeps the elements for which the predicate returns true. Our helper distinctByKey() from the first example takes a key extractor, which is a function that returns the key of an element, for example Song::title. The helper has only two lines, one that creates the seen set and one that returns the lambda.
The predicate relies on Set.add(), which returns true when the key is new and false when the set already holds an equal key. So the first song with a key passes the filter, and every later song with that key is dropped.
The method ConcurrentHashMap.newKeySet() creates a thread-safe Set (that is, several threads can use it at the same time) backed by a ConcurrentHashMap. So the set does not get corrupted when the stream runs in parallel.
For several fields, the key extractor returns a List of the field values. Two lists are equal when they hold equal elements in the same order, so List.of(“rain”, “anna”) is equal to every other List.of(“rain”, “anna”).
List<Song> distinctSongs = songs.stream()
.filter(distinctByKey(s -> List.of(s.title(), s.artist())))
.toList(); // [rain/anna/120, home/ben/90, home/carl/60]
List<Song> byTitle = songs.stream()
.filter(distinctByKey(Song::title))
.toList(); // [rain/anna/120, home/ben/90]
The diagram follows each song through the predicate, and we can see that a song passes the filter only when its key goes into the seen set for the first time.

We can also write a version that takes any number of key extractors. The … in the parameter list (called varargs) lets the caller pass one, two or more functions, and the method builds the List key from them, so the caller lists the fields as method references.
@SafeVarargs
public static <T> Predicate<T> distinctByKeys(Function<? super T, ?>... keyExtractors) {
List<Function<? super T, ?>> extractors = List.of(keyExtractors);
Set<List<?>> seen = ConcurrentHashMap.newKeySet();
return t -> seen.add(extractors.stream().map(ke -> ke.apply(t)).toList());
}
List<Song> distinctSongs = songs.stream()
.filter(distinctByKeys(Song::title, Song::artist))
.toList(); // [rain/anna/120, home/ben/90, home/carl/60]
The predicate keeps the first element for each key, in the original order of a sequential stream. A sequential stream is a normal stream that runs on one thread, and the original order of its elements is called the encounter order.
3. Finding Distinct by Multiple Fields using a Custom Key Class
Instead of a List, the key can be a small class of its own, and a record fits well here because the compiler writes equals() and hashCode() from the record components (its fields). A second constructor builds the key from a Song, and toString() keeps the printed keys short.
record SongKey(String title, String artist) {
SongKey(Song song) {
this(song.title(), song.artist());
}
@Override
public String toString() { return title + "/" + artist; }
}
List<Song> distinctSongs = songs.stream()
.filter(distinctByKey(SongKey::new))
.toList(); // [rain/anna/120, home/ben/90, home/carl/60]
Compared with List.of(…), a key record has three advantages.
- The field names show what makes two songs duplicates.
- The compiler checks the field types, whereas a List<Object> key accepts any value.
- The record accepts null fields, whereas List.of() throws NullPointerException for them (see section 9.2).
We use a key record when the key is used in more than one place, and List.of() for a one-off stream. A key class must be immutable, which means its fields never change after the object is created, because a good HashMap key returns the same hash code for its whole life.
4. Keeping the First or Last Occurrence with toMap()
A shopping app that merges two price lists has the same product twice and must decide which price wins, the old one, the new one or the lower one. A filter cannot make that choice, because it never sees both duplicates at once.

The predicate always keeps the first element, whereas Collectors.toMap() lets us choose which element stays. The collector toMap() puts each song into a map under its key, and when two songs have the same key, it calls the third argument, a merge function, which gets the stored song a and the new song b and returns the one to keep. The fourth argument, LinkedHashMap::new, creates a map that keeps the keys in insertion order. At the end, collectingAndThen() turns the map values into a List.
List<Song> keepFirst = songs.stream()
.collect(Collectors.collectingAndThen(
Collectors.toMap(SongKey::new, Function.identity(), (a, b) -> a, LinkedHashMap::new),
m -> new ArrayList<>(m.values()))); // [rain/anna/120, home/ben/90, home/carl/60]
List<Song> keepLast = songs.stream()
.collect(Collectors.collectingAndThen(
Collectors.toMap(SongKey::new, Function.identity(), (a, b) -> b, LinkedHashMap::new),
m -> new ArrayList<>(m.values()))); // [rain/anna/75, home/ben/40, home/carl/60]
List<Song> keepMostPlayed = songs.stream()
.collect(Collectors.collectingAndThen(
Collectors.toMap(SongKey::new, Function.identity(),
BinaryOperator.maxBy(Comparator.comparingInt(Song::plays)), LinkedHashMap::new),
m -> new ArrayList<>(m.values()))); // [rain/anna/120, home/ben/90, home/carl/60]
Only the merge function changes between the three calls, so the merge function alone decides which song is kept.
| Merge function | Kept element | Result |
|---|---|---|
| (a, b) -> a | First occurrence | [rain/anna/120, home/ben/90, home/carl/60] |
| (a, b) -> b | Last occurrence | [rain/anna/75, home/ben/40, home/carl/60] |
| BinaryOperator.maxBy(comparingInt(Song::plays)) | Highest play count | [rain/anna/120, home/ben/90, home/carl/60] |
In keepLast, the list holds the last song for each key, but the order still follows where each key first appeared, so rain/anna/75 comes first because the key rain/anna first appeared at position 1. Without LinkedHashMap::new, toMap() returns a HashMap, and the result has no fixed order. We also need the merge function here, because toMap() without a merge function throws IllegalStateException at the first duplicate key.
5. Distinct and Sorted with TreeSet and a Comparator
A TreeSet is a set that keeps its elements sorted using a Comparator, and when the comparator returns 0, the TreeSet treats the two elements as duplicates. So when we collect the songs into a TreeSet with a comparator on title and artist, that one step removes duplicates and sorts the result. The comparator calls thenComparing() to sort on multiple fields.
Comparator<Song> byTitleAndArtist = Comparator.comparing(Song::title).thenComparing(Song::artist);
List<Song> sorted = songs.stream()
.collect(Collectors.collectingAndThen(
Collectors.toCollection(() -> new TreeSet<>(byTitleAndArtist)),
ArrayList::new)); // [home/ben/90, home/carl/60, rain/anna/120]
The TreeSet keeps the first element for each key, but the result is sorted by title and artist, not in the original order. We use the TreeSet approach only when we want a sorted result. Sorting also costs more time, because the TreeSet takes O(n log n) while the hash-based approaches take O(n).
6. Distinct by Key in Parallel Streams
Most predicates look only at the current element, but the answer of distinctByKey() depends on the elements it has seen before, which makes it a stateful predicate. For example, a music app that removes duplicate songs from a 100,000-song library in a parallel stream can keep a different copy of each song on every run.
In a parallel stream, several threads work on parts of the list at the same time, and the results of stateful lambdas can change from run to run. The set from ConcurrentHashMap.newKeySet() keeps the set thread-safe, but nothing makes the thread that holds the first element add its key first.
To test the difference, we build 100,000 songs with 100 distinct keys and run both approaches 50 times each in a parallel stream on JDK 25 with two CPU cores. The sequential stream keeps song0/anna/0, the first song of its key.
List<Song> big = IntStream.range(0, 100_000).mapToObj(i -> new Song("song" + i % 100, "anna", i)).toList();
List<Song> sequential = big.stream().filter(distinctByKey(SongKey::new)).toList();
Song firstKept = sequential.get(0); // song0/anna/0
List<Song> parallel = big.parallelStream().filter(distinctByKey(SongKey::new)).toList();
int keys = parallel.size(); // 100
parallel predicate differs from sequential: 50/50
parallel toMap differs from sequential: 0/50
The predicate still returns one song per key, but not the first song per key. The distinct() operation and toMap() with (a, b) -> a do keep the first element in an ordered parallel stream, because each thread works on its own part of the list and the stream combines the partial results in the original order. For parallel streams, we use toMap() with a merge function, or call sequential() before the filter().
7. Distinct by Key with Stream Gatherers (Java 24+)
Java 24 made Stream Gatherers a final feature (JEP 485). A Gatherer is a custom intermediate operation (a step in the middle of a stream, like filter() or map()), and we pass it to stream.gather(). The predicate keeps one seen set for its whole life, whereas a gatherer creates a fresh state for every stream run, so we can reuse one distinctBy() gatherer.
static <T, K> Gatherer<T, ?, T> distinctBy(Function<? super T, ? extends K> keyExtractor) {
return Gatherer.<T, Set<K>, T>ofSequential(
HashSet::new, // new state for each stream run
Gatherer.Integrator.ofGreedy((seen, song, downstream) ->
!seen.add(keyExtractor.apply(song)) || downstream.push(song)));
}
Gatherer<Song, ?, Song> byTitleAndArtist = distinctBy(s -> List.of(s.title(), s.artist()));
List<Song> firstRun = songs.stream().gather(byTitleAndArtist).toList(); // [rain/anna/120, home/ben/90, home/carl/60]
List<Song> secondRun = songs.stream().gather(byTitleAndArtist).toList(); // [rain/anna/120, home/ben/90, home/carl/60]
List<Song> parallelRun = songs.parallelStream().gather(byTitleAndArtist).toList(); // [rain/anna/120, home/ben/90, home/carl/60]
The integrator is the function that handles each element, and it returns true to ask for the next element. For a repeated key, the integrator returns true at once and skips the song, whereas for a new key, it pushes the song downstream to the next step of the stream. The factory method ofSequential() runs the gatherer on one thread, even in a parallel stream, so the first occurrence always wins and a plain HashSet is enough.
The Gatherer API needs Java 24 or later, because Java 22 and 23 had it only as a preview feature. On Java 21 and older releases, the import of java.util.stream.Gatherer fails with “cannot find symbol”, so we use the predicate or toMap() there.
8. Which Approach Keeps Which Element?
All approaches remove duplicates by the same key. They differ in which element they keep and in the order of the result, and the parallel stream column shows which ones keep the first element when several threads run.
| Approach | Kept element | Result order | Parallel stream | Java |
|---|---|---|---|---|
| distinct() | First | Encounter order | First element kept | 8+ |
| filter(distinctByKey(…)) | First (sequential only) | Encounter order | Any element per key | 8+ |
| toMap(…, merge, LinkedHashMap::new) | First, last or by rule | Order of first appearance | Same as sequential | 8+ |
| TreeSet with Comparator | First | Sorted by the comparator | First element kept | 8+ |
| gather(distinctBy(…)) | First | Encounter order | Same as sequential | 24+ |
For most code, filter(distinctByKey(…)) with a record key is the shortest choice. When the last or the best element must win, we use toMap().
9. Stream Distinct by Multiple Fields FAQs
A reused predicate, null fields, String keys, duplicate counts and key-only results cause most of the trouble with distinctByKey() in real code.
9.1. Why Does a Reused distinctByKey() Predicate Return an Empty Result?
The seen set lives inside the predicate object, so a second stream that uses the same predicate object starts with all the keys from the first run. Every element looks like a duplicate in the second run, and the stream returns an empty result.
Predicate<Song> once = distinctByKey(SongKey::new);
long firstCount = songs.stream().filter(once).count(); // 3
long secondCount = songs.stream().filter(once).count(); // 0
So we call distinctByKey() inside each stream pipeline, which gives each stream a new set. The Gatherer version from section 7 creates a new set for every stream run, so we can reuse one gatherer.
9.2. What Happens When a Key Field Is null?
The factory method List.of() does not allow null elements, so a List.of() key throws NullPointerException as soon as a field is null. Arrays.asList() and a key record both accept null.
List<Song> withNull = List.of(new Song("rain", null, 10), new Song("rain", null, 20));
List<Song> listKey = withNull.stream().filter(distinctByKey(s -> List.of(s.title(), s.artist()))).toList(); // NullPointerException
List<Song> arrayKey = withNull.stream().filter(distinctByKey(s -> Arrays.asList(s.title(), s.artist()))).toList(); // [rain/null/10]
List<Song> recordKey = withNull.stream().filter(distinctByKey(SongKey::new)).toList(); // [rain/null/10]
9.3. Can We Concatenate the Fields into a String Key?
A key such as title + artist is short, but two different songs can end up with the same key. For example, “ab” + “c” and “a” + “bc” both give “abc”.
List<Song> tricky = List.of(new Song("ab", "c", 1), new Song("a", "bc", 2));
List<Song> byString = tricky.stream().filter(distinctByKey(s -> s.title() + s.artist())).toList(); // [ab/c/1]
List<Song> byRecord = tricky.stream().filter(distinctByKey(SongKey::new)).toList(); // [ab/c/1, a/bc/2]
A separator between the fields lowers the risk, but the risk stays when the separator can appear in the data, so we prefer a List or record key.
9.4. How Do We Count Distinct Elements or the Duplicates per Key?
The terminal operation count() after the predicate returns the number of distinct keys. The collector groupingBy() with counting() returns how often each key appears, and with mapping() it shows which values were merged under each key. The same collectors help us find and remove duplicates by any key.
long distinctCount = songs.stream().filter(distinctByKey(SongKey::new)).count(); // 3
Map<SongKey, Long> countPerKey = songs.stream().collect(Collectors.groupingBy(SongKey::new, LinkedHashMap::new, Collectors.counting()));
// {rain/anna=3, home/ben=2, home/carl=1}
Map<SongKey, List<Integer>> playsPerKey = songs.stream().collect(Collectors.groupingBy(SongKey::new, LinkedHashMap::new,
Collectors.mapping(Song::plays, Collectors.toList())));
// {rain/anna=[120, 75, 75], home/ben=[90, 40], home/carl=[60]}
9.5. How Do We Get Only the Distinct Key Values?
Sometimes we want only the field values, not the whole songs. In that case, we map each song to its key and call distinct(), and the equals() method of the key record decides which keys are equal.
List<SongKey> keys = songs.stream().map(SongKey::new).distinct().toList();
// [rain/anna, home/ben, home/carl]
10. Conclusion
Java streams have no built-in distinctBy() method, so we write a short distinctByKey() predicate that stores List or record keys in ConcurrentHashMap.newKeySet(). The predicate removes duplicates by any mix of fields and keeps the first element in encounter order.
Collectors.toMap() with a merge function and LinkedHashMap::new can keep the first, the last or the best element, and it stays correct in parallel streams. A TreeSet with a Comparator fits when the result should be sorted, and on Java 24+, a Gatherer gives us a reusable distinctBy() operation. Finally, we create a new predicate for each stream and avoid List.of() keys when a field can be null.
11. References
- Stream Javadoc (Java 25)
- Collectors Javadoc (Java 25)
- ConcurrentHashMap Javadoc (Java 25)
- TreeSet Javadoc (Java 25)
- Gatherer Javadoc (Java 25)
- JEP 485: Stream Gatherers
Happy Learning !!
How can we get same list of Record with all distinct count as comma seperated? consider count is String field
Record [id=1, count=[10,20], name=record1, email=record1@email.com, location=India],
I noticed that both of the above solutions will allocate a new object on every single element of the stream. For a large stream, that will be a lot of garbage created that needs to be collected. Here is a solution I came up with that avoids allocation on every single element of the stream. Seems to work welll:
public final class StreamHelpers { private StreamHelpers() { } @SuppressWarnings("unchecked") public static <T> Function<T, Stream<T>> distinctBy(Function<T, ?>... fieldSelectors) { ConcurrentMap keysSeen = new ConcurrentHashMap<>(); return t -> { ConcurrentMap keyMap = keysSeen; for (int i = 0; i < fieldSelectors.length - 1; i++) { Function<T, ?> selector = fieldSelectors[i]; keyMap = (ConcurrentMap) keyMap.computeIfAbsent ( selector.apply(t), k -> new ConcurrentHashMap() ); } boolean seen = (boolean) keyMap.compute ( fieldSelectors[fieldSelectors.length - 1].apply(t), (k, v) -> v != null ); return seen ? null : Stream.of(t); }; } }It is used like:
Unfortunately, it still allocates once for each element the first time that element’s key is seen (Stream.of(…)). However, if there are a significant number of duplicates, this should reduce the allocations significantly. It would be nice to have a solution that involved zero allocations if possible though. I made it work using flatMap instead of filter because I felt having a Predicate that has side-effects violates the contract of filter. I was trying to get something that follows all the contracts of the stream methods AND avoids all allocations (or at least minimizes them). Any ideas for how to get there better?
Hello ,
i followed your tutorial , right now it showing my desired output, but still i need to get the total count of the filtered item and put it in object list -> so i can show how many count from each group
Hope you can help me , Thanks