Inspiration
Pitt's Wireless & Telecom Services runs nine residence-hall mailrooms. The brief was straightforward: package volume spikes at move-in and around holidays, staff get overwhelmed, so predict the surges and we'll use it to justify buying more lockers.
We built that forecast. Then the data told us the premise was wrong.
What it does
Panther Post forecasts daily parcel intake per mailroom, then derives everything operations actually needs from it:
Occupancy — intake convolved with a dwell survival curve → how many parcels are physically sitting there Pickups — intake convolved with the dwell hazard → counter workload Staffing — (intake + pickups + disposals) × measured handling time → people to roster, day by day
One model per mailroom, six in total. Output is a week plan, a Power BI dashboard, and a spoken briefing.
Every day is published at two levels: the median, and the 80th percentile. Staffing plans against q80, because a short shift costs more than an idle one.
How we built it
Two years of real WTS data — 109,482 parcels, 2024-09-19 to 2026-09-19. Entirely classical statistics and machine learning; no language models anywhere in the pipeline.
Six estimator families (Poisson, Poisson+L2, Tweedie, Bayesian ridge, gradient-boosted Poisson trees, hierarchical) crossed with four feature sets Selection by rolling-origin cross-validation inside the training year, scored on pinball loss at q80 so model choice matches the asymmetric staffing cost Discrete Kaplan-Meier survival estimation for dwell, censoring administrative purges and still-open records so they aren't miscounted as student pickups Year one trains, year two tests and is never touched during selection Per-parcel handling times measured from gaps between consecutive scans (37–49s by site), so no time-motion study was needed Challenges we ran into
Arrivals are not occupancy. Tower B's busiest day takes in 741 parcels. Its peak concurrent holdings are 3,596 — nearly five times that. Median dwell is about a day, but 22% of parcels are still held after a week and 7% after a month, so holdings accumulate far above any single day's intake.
A log-link model fed its own predictions diverges. Forecasting 36 days out with rolling-average features drove Tower B's 7-day mean from 127 to 188,400 and the Poisson mean to infinity. Fixed with horizon-aware selection: momentum features within 14 days, calendar-only beyond.
Detecting an administrative purge needs age, not volume. A volume test flags every post-break resumption rush. Keying on median parcel age separates a bulk disposal (2026-07-01: 1,806 items, median age 210 days) from students returning in January (419 items, median age 35 days).
One site changed behaviour and no other did. Tower B's share of parcels held past 90 days went 0.9% → 6.7% between academic years while every other site stayed under 0.5%. One pooled survival curve couldn't represent both regimes; stratifying cut occupancy error from 44.6% to 13.7%.
Data hygiene under deadline. We'd committed a live API key and 46 MB of data files to a public repo. Rotating the key and rewriting history cost hours we'd rather have spent modelling.
Accomplishments that we're proud of
The forecast overturned the question. Tower B's capacity problem is non-collection, not demand. A 30-day return-to-sender policy cuts peak holdings from 2,223 to 227 parcels — 90% — and costs nothing.
All six models clear the bar. Every site beats seasonal naive (last week, same weekday) by 11.7% to 41.2% on a test year the models never saw.
Occupancy is validated, not asserted. The export carries both received and delivered timestamps, so true occupancy is directly observable and the dwell model gets a free validation set. Tower B's occupancy error is 14.3% of mean holdings.
The recommendation costs nothing. Peak staffing is 7.5 FTE against a 9.5 pool — an assignment from existing staff, not a hiring ask.
What we learned
Measure what you actually care about. We spent our first day predicting arrivals because that's what was asked. Occupancy was sitting in the same export the whole time.
Pick the metric that matches the cost. R² ranked models misleadingly at single-digit daily counts — our earlier version reported R² of −0.06 for a model beating its baseline by 57% on MAE. Pinball loss at q80 reflects what understaffing actually costs.
Verify your own claims. Our README stated the export contained recipient names. Audited column by column, it doesn't — that field holds carrier codes. We had propagated an assumption nobody checked.
.gitignore is not retroactive, and a force push doesn't remove orphaned commits from GitHub.
What's next for PackageSurgePredictor Real locker counts from Operations. The 1,000-unit pool is a placeholder; it's the only reason Bigelow reads 121% utilisation. Ruskin. Opened 2025-09-09, so it has no training-year data and is skipped rather than guessed at. A hierarchical model predicting it as a share of Tower B would close the gap. Saturdays. q80 coverage 0.753 against a 0.80 target — the weakest weekday. Take the disposal finding to WTS. It needs a policy decision, not more modelling. Nightly refresh — predict → make_powerbi → briefing on a schedule, feeding the dashboard and pre-generating the audio.
Log in or sign up for Devpost to join the conversation.