SELECTED WORK

Things I've built around real-world data.

A selection of machine learning, AI, and data engineering work from academic projects and professional experience.

3 SELECTED SYSTEMS

Property Price Prediction

Applied machine learning for residential property pricing

Can spatial and neighborhood-derived features improve property price prediction?

0.8882 R² SCORE
≈ $23K MAE
11 MIAMI CITIES

Problem

Predicting property prices requires more than basic property attributes. Location and neighborhood effects can strongly influence the final value.

Approach

Built a predictive modeling workflow using XGBoost and CatBoost, with Scikit-learn for preprocessing and evaluation.

Features included

  • Spatial clustering
  • KNN average price
  • Naive Bayes ZIP-code probability
  • Property attributes
  • Location-based features
RAW PROPERTY DATA
↓
FEATURE ENGINEERING
↓
SPATIAL FEATURES
↓
MODEL
↓
PRICE PREDICTION
Hover a node to see details.
Python · XGBoost · CatBoost · Scikit-learn · KMeans · KNN · FastAPI

AI Property Description & Summary Generation

Turning structured property data into useful natural-language content

How can structured property data be transformed into consistent, useful language?

Built an LLM-powered content pipeline that aggregates building and listing information from Elasticsearch and generates structured building descriptions and buyer-oriented AI summaries.

ELASTICSEARCH
↓
BUILDING + LISTING CONTEXT
↓
CONTEXT AGGREGATION
↓
OPENAI / GEMINI
↓
STRUCTURED JSON
↓
APPLICATION OUTPUT
Hover a node to see details.
Python · OpenAI · Gemini · Elasticsearch · Pydantic

MLS Data Pipeline & Entity Resolution

Matching and attributing property listings at scale

How can large volumes of listing data be cleaned, matched, and attributed reliably?

113M+
MLS PROPERTY RECORDS
~4TB
PRODUCTION DATA

Built scalable Python and Elasticsearch workflows for extracting, cleaning, matching, and updating MLS listings across the US.

GEOSPATIAL
+
POLYGON
+
PRICE
+
TEXT
↓

LISTING ATTRIBUTION
Hover a node to see details.
Python / Elasticsearch bulk processing
~598K documents per cycle

What these projects taught me.

DATA

Working with real-world data means dealing with quality, ambiguity, and scale.

MODELS

Good models depend on meaningful features, sound evaluation, and understanding the problem.

SYSTEMS

A useful solution needs reliable engineering around the model.