Sachin Goel created FLINK-2131:
----------------------------------
Summary: Add Initialization schemes for K-means clustering
Key: FLINK-2131
URL:
https://issues.apache.org/jira/browse/FLINK-2131 Project: Flink
Issue Type: Task
Components: Machine Learning Library
Reporter: Sachin Goel
Assignee: Sachin Goel
The Lloyd's [KMeans] algorithm takes initial centroids as its input. However, in case the user doesn't provide the initial centers, they may ask for a particular initialization scheme to be followed. The most commonly used are these:
1. Random initialization: Self-explanatory
2. kmeans++ initialization:
http://ilpubs.stanford.edu:8090/778/1/2006-13.pdf3. kmeans|| :
http://theory.stanford.edu/~sergei/papers/vldb12-kmpar.pdfFor very large data sets, or for large values of k, the kmeans|| method is preferred as it provides the same approximation guarantees as kmeans++ and requires lesser number of passes over the input data.
--
This message was sent by Atlassian JIRA
(v6.3.4#6332)