There’s a sealed VHS copy of Akira averaging $1,544 on eBay right now. The Matrix on LaserDisc sits at $1,098. And the Coen Brothers’ Ballad of Buster Scruggs — a Netflix original that barely got a physical release — averages $1,293 on Blu-ray. Five years ago, VHS tapes were garage sale filler at fifty cents each. Now some of them fetch four figures. But nobody has a systematic way to track it. No price history. No trend analysis. No signal on what’s heating up before it peaks. I’m a collector. Cards, comics, coins, watches — I’ve spent decades watching how niche markets move. And the film collectibles market right now looks exactly like trading cards did in the mid-2010s: exploding in value, driven by nostalgia and scarcity, with zero infrastructure for understanding price dynamics. So I built the infrastructure. Film Price Guide (filmpriceguide.online) is an automated pipeline that crawls eBay twice daily, monitors 16 film subreddits for demand signals, runs ML predictions on price movements, and serves it all through a searchable web interface. Seven months of continuous operation. 901,713 price records. 3,770 movies tracked across DVD, Blu-ray, 4K, VHS, and LaserDisc. This is how I built it and what the data actually tells you. The Architecture The system is a daily pipeline running on a single Linux server — no cloud dependencies, no managed services, just cron and discipline. The stack: PHP 8.4 handles the web app and API crawling (Guzzle is excellent for eBay’s API). Python 3.12 handles the ML and analysis work (pandas, scikit-learn). PostgreSQL 15 stores everything. Redis handles caching. Ollama runs a local LLM (llama3.1:8b) for market narrative generation. Alpine.js on the frontend. Two languages in one pipeline sounds messy. It’s actually clean — a shell script orchestrates everything, calling each interpreter directly. PHP does what PHP does well (HTTP, web rendering). Python does what Python does well (data science). Trying to force everything into one language would have been the fragile choice. Phase 1: Data Collection The core is a PHP-based eBay crawler that fires on cron at 2 AM and 2 PM. Each run picks 100 movies from a TMDB-sourced catalog of 3,770 films, searches eBay’s API across all five formats, and stores the results as JSONB in PostgreSQL. JSONB was deliberate. eBay listings have wildly inconsistent fields — condition descriptions, grading metadata, shipping details — and trying to normalize that into rigid columns would have been a maintenance nightmare. JSONB lets me store the full payload while still querying specific fields with indexes. Worth the storage trade-off. After seven months: 901,713 total price records 3,753 movies with pricing data (out of 3,770 in the catalog) ~7,500 new listings collected daily Here’s where the data gets interesting. The format breakdown tells you the shape of the physical media market: DVD is the high-volume, low-margin format. Blu-ray is the mainstream sweet spot. 4K is where enthusiasts play. But LaserDisc — a dead format, discontinued in 2009 — commands the highest average at $42.05. That’s double Blu-ray and triple DVD. Low supply, dedicated collector base, almost no new inventory entering the market. Classic scarcity dynamics. VHS punches above its weight at $22.03, but that average is pulled up hard by sealed and graded specimens. The Akira VHS at $1,544 isn’t a typical VHS sale — it’s a graded collectible that happens to be a VHS tape. Phase 2: Wave Analysis Raw pricing data is table stakes. The real question: when should you buy, and when should you sell? I built a wave analysis system that combines three signal sources into a single score from 0 to 100. Signal 1: eBay Scarcity (40% weight) For each movie, I calculate listing count, sell-through rate, and a scarcity rating ranging from “abundant” to “extremely rare.” A movie with 5 active listings and 80% sell-through rates much higher than one with 200 listings and 10% sell-through. This is the fundamental supply/demand signal. Signal 2: Reddit Sentiment (40% weight) I scan 16 subreddits daily — everything from r/VHS and r/criterion to r/flipping and r/boutiquebluray. For each post, I extract movie title mentions (regex plus database matching against the full 3,770-movie catalog), engagement metrics, cross-subreddit spread, and price mentions. The trending score formula weights mentions, upvotes, comments, and subreddit breadth. A movie mentioned across five different collecting communities is a stronger signal than one with a single viral post. In the last 30 days, 120 movies had measurable Reddit signal. Top performers: Spirited Away (wave score 66.3), Vertigo (78.9 — high scarcity, low listings), The Insider (78.6), Pulp Fiction (64.0), The Thing (64.0). Signal 3: Seasonal Patterns (20% weight) Seven months of data confirmed what collectors intuitively know: winter is selling season. High volume, lower prices — sellers dumping holiday inventory. Summer commands a premium. I apply seasonal multipliers: summer 1.3x, spring 1.1x, fall 1.0x, winter 0.85x. Early versions hard-coded these adjustments. Now they’re a feature the ML model can learn to weight or ignore based on actual data. That shift from rule to feature was a small change with big implications for model flexibility. What the Wave Scores Tell You After 38,075 wave analyses, the distribution looks like this: 82% sit at “Bottom” (score under 25) — quiet, stable, low activity 13% are “Cooling down” 5% are “Stable” Less than 1% are “Heating up” Less than 0.1% are at “Peak” This matches every niche market I’ve ever watched. The vast majority of inventory sits quietly. The opportunity is in catching the small percentage that’s gaining momentum before it peaks. That’s the whole point of the system. Phase 3: Machine Learning Wave scores are rule-based. I wanted to go further — can a model actually predict whether prices will go up or down? The Setup I joined wave analysis features with actual 7-day forward price changes from the price_history table. That gave me 7,868 labeled samples — each one a movie at a point in time with known features and a known outcome. Twelve features per sample: Reddit mention count, trending score, and sentiment; eBay listing count, sold count, active count, and average price; scarcity score; seasonal multiplier; encoded season and format; and graded count. I trained both RandomForest and GradientBoosting regressors and auto-selected the better performer. The training script runs weekly on Sunday at 1 AM. The Honest Results First model training: Model: RandomForest (200 estimators, max_depth=10)
Target: 7-day price change (%)
Train R²: 0.2782
Test R²: 0.0099
Test MAE: 19.15% An R² of 0.01 on test data. I’m not going to pretend that’s good. It’s not. Predicting individual movie price swings over a 7-day window in a thin, sentiment-driven market is inherently noisy. If someone tells you they can do this with high accuracy on sparse data, they’re either overfitting or lying. But here’s why it’s still useful. The feature importance breakdown: eBay supply/demand features carry over 90% of the weight. Reddit signal contributes almost nothing — because most movies have zero mentions in any given period. The signal is too sparse right now. But as data accumulates and the subreddit coverage expands, that should change. The model retrains weekly, so improvement is structural. The Blend Rather than replacing the rule-based wave analysis, the ML model enhances it: If ML agrees with the rule-based prediction, confidence gets boosted. If ML disagrees and model R² is above 0.3, trust ML more (60/40 blend). If ML disagrees but the model is still weak, keep the rules and lower confidence. This gives me a graceful degradation path. The rules work now. The model gets better over time. The system doesn’t break during the transition. Phase 4: Making Reddit Actually Useful The original pipeline analyzed 300 random movies per run. Most had zero Reddit signal. That’s a lot of noise and not much pattern. I redesigned it to be signal-first: query all movies with Reddit activity in the last 7 days and analyze those first, then fill remaining slots with random selections for broad coverage. Instead of 300 random movies with no Reddit mentions, I now get 25+ Reddit-active movies with real engagement data plus ~275 random movies for baseline coverage. The Reddit scraper itself got upgraded: 16 subreddits (up from 9), three sort modes per subreddit (hot, top, and new — not just hot and top), and database-level title matching that catches quoted and markdown-formatted titles even without a year attached. The single biggest quality improvement in the whole project wasn’t a model tweak — it was prioritizing signal over random sampling. A wave analysis of a movie nobody’s talking about is noise. A wave analysis of Spirited Away during a Reddit hype cycle is signal. What I Learned From 900,000 Records LaserDisc is the sleeper format. Average $42.05 per listing, double Blu-ray. Dead format, dedicated base, no new supply. If you’re looking for an undervalued corner of the market, start here. Scarcity beats fame. Buster Scruggs at $1,293 outprices The Godfather in every format. Limited physical releases of streaming-first titles create unexpected premiums. I’ve watched this exact pattern play out in sports cards (short-print variants), comics (limited covers), and coins (low-mintage years). The mechanism is universal — the market is just catching up in film. 82% of the market is quiet. Most movies sit at bottom-level pricing with stable-to-low activity. The entire game is catching the 5% that are heating up before they peak. Which is exactly why the wave system exists. Reddit is a leading indicator — when it fires. Movies trending on r/criterion and r/boutiquebluray see price movement within one to two weeks. But only 120 movies had any Reddit signal in the last 30 days. It’s a narrow funnel. Expanding subreddit coverage and adding comment-level analysis (not just post titles) is the next priority. Winter is for buying. Highest volume, lowest prices. Spring and summer climb as convention season and nostalgia kick in. If you’re building a collection, December through February is your window. What’s Next Heritage Auctions integration — premium auction data is already in the database but not yet feeding the model. Heritage lots of graded VHS regularly clear $10K+, and that price ceiling data should improve predictions at the high end. Price alerts — the webhook infrastructure is in place. Next step: let users set target prices and get notified. Comment-level Reddit analysis — post titles are a blunt signal. The real demand indicators live in comment threads where people are asking “where can I find this?” and “how much did you pay?” And the model keeps retraining. Every week, more labeled outcomes. Every week, slightly less noise and slightly more signal. That’s how you build a prediction system that actually improves — not by chasing accuracy on day one, but by building the pipeline that makes accuracy inevitable over time. The Numbers Built with PHP, Python, PostgreSQL, scikit-learn, and the conviction that every niche market deserves better data infrastructure than a gut feeling and an eBay search. I write about applied ML, pattern recognition in messy domains, and building systems that find signal in noise. More at theorubin.com.