Module Question 5
- Describe the steps of the k-Means clustering algorithm. What does the ‘k’ represent?
- Why is dimensionality reduction useful in machine learning?
- Provide a real-world application for clustering (e.g., customer segmentation).
- What is the main objective of PCA?
Status: 100% sudah tercapai
Keterangan: saya sudah mengerjakan dengan baik dan benar
Bukti:
Answer:
1.k-Means is an unsupervised machine learning algorithm used to group similar data points into clusters.
Steps:
-Choose the number of clusters, k.
-Randomly select k points as the initial cluster centers (centroids).
-Assign each data point to the nearest centroid.
-Recalculate the centroid of each cluster.
-Repeat the assignment and recalculation steps until the centroids stop changing significantly.
‘k’ represents the number of clusters that we want to create.
2.Dimensionality reduction reduces the number of features (dimensions) in a dataset while keeping the most important information.
It is useful because it:
- Makes data easier and faster to process.
- Reduces computational complexity.
- Can reduce noise and irrelevant information.
- Helps prevent overfitting.
- Makes high-dimensional data easier to visualize.
3.A common application is customer segmentation.
For example, a company can use customer data such as:
- Age
- Income
- Purchase frequency
- Amount spent
A clustering algorithm can group customers into categories such as:
- High-value frequent customers
- Occasional customers
- Price-sensitive customers
The company can then create different marketing strategies for each group.
4.PCA (Principal Component Analysis) aims to reduce the number of dimensions in a dataset while preserving as much of the important variation/information as possible.
In simple terms, PCA transforms many related features into a smaller number of new features called principal components.
Example:
If a dataset has 10 features, PCA might reduce them to 2 or 3 principal components while retaining most of the useful information.
