How PuckAI is built
PuckAI is a solo-developed hockey analytics platform: NHL player profiles, prospect projections, cross-era comparisons and an AI scouting chat. The design splits cleanly in two — a batch pipeline that computes everything into reproducible, versioned artifacts, and a stateless read API that serves them. The full write-up, with the decision records and the measurements behind every number here, is public.
Read the full write-upCompute offline, serve online
Sources
NHL API · EliteProspects · Central Scouting
14-stage ETL
clean → filter → features → tiers → export
XGBoost heads
entry · longevity · tier
Reports + validation
generated, then checked against the pipeline
Embeddings
1024-d vectors over every report
Import scripts
versioned artifacts loaded into Postgres
PostgreSQL 16 + pgvector
HNSW cosine index, hosted on Supabase
FastAPI
read-only services, serverless
Next.js 16
App Router, React 19
Hockey data is read-only online: nothing in the request path recomputes a tier, a projection or an embedding. The only online writes are user-owned — chat conversations, feedback and saves.
What the pipeline produces
Player Tiers
A five-level classification — Elite, Star, Core, Depth, Fringe — over 2,700 NHL skaters, with eight-dimensional aspect scoring for every season.
Prospect Projection
Three XGBoost models predicting NHL entry likelihood, career longevity (≥ 200 games) and tier placement across 55,455 pre-NHL players.
Similar Player Search
pgvector nearest-neighbour retrieval over 8-dimensional aspect vectors, supporting cross-era comparisons from 1997-98 onward.
Scouting Reports
7,643 algorithmically generated reports, each validated against pipeline data before publication and carrying full source attribution.
Scout Chat
Retrieval-augmented search combining Postgres full-text indexing, dense vectors, weighted reciprocal rank fusion, neural reranking and scope routing — answers arrive with citations.
Measured, not asserted
Retrieval quality
77 labeled queries, NDCG at depth 10.
- Keyword only0.488
- Vector only0.508
- Hybrid + rerank0.582
- Hybrid + scope routing0.682
Prospect prediction
- P(NHL)
- 0.977 ± 0.001 ROC-AUC · 0.849 ± 0.005 PR-AUC
- P(GP ≥ 200 | NHL)
- 0.699 ± 0.014 ROC-AUC
- Tier — Forward
- 0.375 exact · 0.767 within-1 · 0.379 macro-F1
- Tier — Defenseman
- 0.314 exact · 0.729 within-1 · 0.291 macro-F1
Report generation
953 batches, July–August 2026.
- Generated
- 7,670 attempts
- Passed validation
- 7,643 (99.65%)
- Rejected & regenerated
- 27 (0.35%)
What is in the database
| Component | Count |
|---|---|
| Pre-NHL seasons | 859,556 |
| Prospect profiles | 55,455 |
| Served predictions | 54,933 |
| NHL season trajectories | 38,899 |
| Scouting reports (1024-d) | 7,643 |
| NHL tier assignments | 6,078 |
| League adjustment factors | 1,484 |
| NHL skaters (full profiles) | 2,700 |
The stack
Offline
- Python 3.12
- pandas
- XGBoost
- scikit-learn
- Staged ETL with a CLI runner
- pytest validation
Online
- FastAPI 0.136
- SQLAlchemy 2.0 · Pydantic 2.13
- PostgreSQL 16 + pgvector (HNSW)
- Supabase
- Next.js 16 · React 19
- TypeScript · Tailwind 4
AI
- Voyage embeddings (1024-d)
- rerank-2.5-lite
- Pluggable answer model
Development
- Docker Compose for the local stack
Read the engineering write-up
The public repository holds documentation only — no product code. It covers:
- Architecture patterns and request flow
- 14-stage ETL guarantees and idempotency
- Hybrid search implementation and result tables
- The report validation checker — six checks, no LLM
- Measurement methodology and its limitations
- Debugging notes: label leakage in Stage 1, age-band calibration
- Design decision records, with the tradeoffs