SOFTWARE ENGINEER & SDET

Jennifer Montgomery

Backend · full-stack · quality engineering

COURSEWORK / TRADE & AHEAD

Unsupervised learning and clustering

Can stocks be grouped by observed price and financial characteristics without a target label?

← Coursework on resume

Completed UT Austin postgraduate coursework using a supplied scenario, dataset, and starter notebook. The charts below come from my completed notebook.

UNSUPERVISED LEARNING

Trade & Ahead

COURSE FOCUS

Scaling, K-means, hierarchical clustering, and cluster interpretation

LIBRARIES USED

pandas, NumPy, scikit-learn, SciPy, Seaborn, Matplotlib

The supplied data

The course supplied 340 company records with ticker, sector, sub-industry, stock price, 13-week price change and volatility, plus financial measures such as return on equity, cash ratio, cash flow, income, earnings per share, shares outstanding, and valuation ratios. There was no “correct cluster” label to predict.

What I did

  • Explored skew, outliers, and sector mix, then prepared and scaled the numeric features for distance-based methods.
  • Compared candidate K-means group counts with elbow, distortion, and silhouette views; also examined hierarchical clustering.
  • Profiled the resulting groups by price, price change, volatility, and financial indicators instead of treating a cluster ID as an investment rating.
Silhouette scores across candidate numbers of K-means stock clusters, highest around two clusters.
From the notebook: the silhouette comparison favored two groups, while the elbow view offered a different tradeoff. Open chart ↗
Two-dimensional t-SNE projection of 340 stocks, colored by their two K-means cluster labels.
From the notebook: a two-dimensional view of the fitted groups. Projection helps inspect separation but does not show all original financial dimensions. Open chart ↗

Finding in the course dataset

The two-cluster solution grouped 307 stocks together and 33 in a smaller group. The smaller group showed lower average price, negative average 13-week price change, and greater volatility in the notebook’s profile. These are descriptive patterns in the supplied sample, not a forecast.

Learning reinforced

This exercise showed how preprocessing and the choice of cluster count shape an unsupervised result, and why group descriptions require looking back at the original variables. The clusters are an analysis aid, not investment advice or a portfolio recommendation.

ALL CASE STUDIESCoursework overview ↗