Production-ready K-Means clustering for Apache Spark with pluggable Bregman divergences (KL, Itakura-Saito, L1, etc). 6 algorithms, 740 tests, cross-version persistence. Drop-in replacement for MLlib with mathematically correct distance functions for probability distributions, spectral data, and count data.
- kullback-leibler-divergence
- clustering
- spark-mllib
- cosine-similarity
- k-means
- itakura-saito-divergence
- bregman-divergence
- euclidean-distance
- embeddings
- spark
- entropy
- similarity-search
Scala versions:
2.10
Found 1 artifact