Production-ready K-Means clustering for Apache Spark with pluggable Bregman divergences (KL, Itakura-Saito, L1, etc). 6 algorithms, 740 tests, cross-version persistence. Drop-in replacement for MLlib with mathematically correct distance functions for probability distributions, spectral data, and count data.
- euclidean-distance
- bregman-divergence
- cosine-similarity
- itakura-saito-divergence
- embeddings
- similarity-search
- k-means
- clustering
- kullback-leibler-divergence
- spark
- entropy
- spark-mllib
Scala versions:
2.10
Found 1 artifact