distinct() 是 stream 的中间操作,基于 equals() 和 hashcode() 去重,保留首次出现元素;自定义类需重写 equals/hashcode,否则按内存地址比较;也可用 tomap 或 filter+concurrenthashmap 按字段去重。

Java 中的 Stream 本身没有叫 distinct() 的“方法”,但有一个 中间操作 distinct(),它能根据元素的 equals() 和 hashCode() 去重,是去除流中重复元素最常用、最简洁的方式。
distinct() 的基本用法
它适用于任何引用类型(如 String、Integer、自定义对象等),前提是对象正确重写了 equals() 和 hashCode()。
- 对基本包装类(
Integer、String等)默认已实现,可直接用 - 对自定义类,若未重写
equals/hashCode,则按内存地址比较,相同对象实例才去重,逻辑上通常不符合预期
示例:
List<string> list = Arrays.asList("a", "b", "a", "c", "b");
List<string> distinctList = list.stream()
.distinct()
.collect(Collectors.toList());
// 结果:["a", "b", "c"]
</string></string>
对自定义对象去重的关键:重写 equals 和 hashCode
比如有一个 User 类,想按 id 去重:
public class User {
private Long id;
private String name;
// 构造、getter 省略
@Override
public boolean equals(Object o) {
if (this == o) return true;
if (o == null || getClass() != o.getClass()) return false;
User user = (User) o;
return Objects.equals(id, user.id); // 只看 id 是否相等
}
@Override
public int hashCode() {
return Objects.hash(id); // 保证 equals 为 true 时 hashCode 也相同
}
}
之后就能正常去重:
List<user> users = Arrays.asList(
new User(1L, "Alice"),
new User(2L, "Bob"),
new User(1L, "Alicia") // id 相同,会被去重
);
List<user> uniqueUsers = users.stream()
.distinct()
.collect(Collectors.toList()); // 只保留第一个 id=1 的 User
</user></user>
按某个字段去重(不改 equals)的替代方案
如果不想修改类的 equals 语义(比如 User 本应所有字段都参与比较),但只想按 id 去重,可以用 Collectors.toMap() 或 filter() + 自定义状态:
- 用
toMap保留首次出现的元素(推荐,清晰且线程安全):
List<user> uniqueById = users.stream()
.collect(Collectors.collectingAndThen(
Collectors.toMap(User::getId, Function.identity(), (a, b) -> a),
map -> new ArrayList(map.values())
));
</user>
- 用
filter+ConcurrentHashMap记录已见 key(适合并行流):
Set<long> seenIds = ConcurrentHashMap.newKeySet();
List<user> unique = users.parallelStream()
.filter(user -> seenIds.add(user.getId()))
.collect(Collectors.toList());
</user></long>
注意点和常见误区
distinct() 是有状态的中间操作,依赖流的数据源顺序 —— 它保留**第一次出现**的元素,后续重复项被跳过。
- 它不能指定“按哪个字段”去重,完全依赖对象自身的
equals/hashCode - 不支持原始类型流(如
IntStream),因为int没有equals;但IntStream本身重复值少,一般先转成装箱流再用distinct() - 性能上,
distinct()内部使用HashSet缓存已见元素,时间复杂度 ~O(n),空间复杂度 ~O(n)
不复杂但容易忽略细节。用对了 distinct(),去重就很简单;用错了,可能一个重复都去不掉。
Java免费学习笔记:立即使用
解锁 Java 大师之旅:从入门到精通的终极指南











