-
-
SideQuests turns your free time into a plan nearby and gets you back on time.
-
Set up once, say what you feel like doing, then join plans, split costs and rate what you did.
-
The bigger goal: a social network that matches travelers to places and to each other.
-
The tech stack: a SwiftUI app, a Go API behind NGINX, MongoDB, a Python ML service and a Muse Spark crawler.
-
The ML stack: a 1.85M-parameter ranker trained on synthetic data beats a tuned baseline by 39% on Spearman.
SideQuests
Turn waiting into wandering. Tell SideQuests where you are, when you need to be back and what you're in the mood for. It builds a timed plan of nearby events and places, with the travel between them, that fits your taste. The plan gets better every time you rate an outing, and you can open it up for people nearby to join.
Note: for information about our HEARSAY submission, please scroll to the bottom!
Inspiration
Three of us (Bayan, Kevin and Karthik) met on a study-abroad program this summer, and Max has plenty of his own travel stories. We kept running into the same problem: free time in a place we didn't know, and no fast way to turn it into something worth doing.
- Bayan in Belgium: they asked what to do and the internet said "eat waffles." They found some overpriced, mediocre waffles, saw a couple of tourist attractions a local suggested, then went straight back to the train station because nothing else came to mind.
- Max's layovers: in France, he and his mom planned a transit-only day with ChatGPT and got a generic list with no timing or routes. In Vienna, he and his dad wanted schnitzel and sightseeing. They took badly ordered routes, reached the sights after dark, and found every kitchen closed.
- Karthik in Berlin: he found that he loves travel when it fits him: nature and small groups in the Tiergarten, plus the city's social events. A generic "top 10" list would never have found that mix.
- All summer: we organized trips in the most decentralized way possible: across a travel app, a calendar, a group chat and two payment apps. We knew there had to be something better.
Google Maps, ChatGPT, Fever, Wanderlog and meetup apps like WashedUp each cover one piece. However, none of them carry the whole package. SideQuests turns events, places and transit into one timed plan, learns from your ratings, and lets others join until the plan locks.
What it does
- Taste onboarding: you rate ten interests from 1 to 5, set your company, pace and spend, and answer three short prompts. You can also import your likes from Facebook, which the app turns into suggested ratings for you to approve. From all of this, we build a likes embedding and a dislikes embedding.
- A sidequest in about 30 seconds: pick where you start and end, how far you'll go, and when you need to be back, then type or dictate your mood ("chill and outside, then live music"). In 0.2–0.6 s you get three timed plans with the walk, transit or rideshare legs between stops. Drag to reorder and the route re-times itself. Press and hold a stop to swap it for something similar.
- Social: keep a plan private, share it with friends, or post it to the Forum with a group-size cap and a lock time. You can also post that you're free right now. Every group gets a chat, a shared photo album and an expense split that's exact to the cent. Updates arrive live over WebSockets.
- Agentic checkout: for ticketed activities, an agent finds tickets on the official site, fills in your details, waits for your approval (or buys instantly under a limit you set) and books them. The ticket then appears on that stop for everyone in the group.
- The feedback loop: you rate each stop after the outing, and those ratings update your taste profile, so your next plan starts smarter.
How we built it
SwiftUI app ──HTTPS/WSS──► Cloudflare ► nginx ► Go API ──► MongoDB (catalog + app data)
│
└──► FastAPI ML service ──► Qwen3 embeddings
classifier · Jev rerank (Vertex → HF → local)
Offline: ingestion crawler (Ticketmaster, Google Places, OSM, RA → Muse → Gemini) → MongoDB
Academic Compute (4× A100): synthetic data + training → Hugging Face Hub → ML service
| Layer | Stack |
|---|---|
| iOS | Swift, SwiftUI, WebSockets, Keychain, Apple Speech, ASWebAuthenticationSession, XCTest/XCUITest |
| Backend | Go, MongoDB (2dsphere, TTL), JWT with rotating refresh tokens, a realtime hub, a checkout agent |
| ML | FastAPI, PyTorch, sentence-transformers, Qwen3-Embedding-0.6B, Qwen3.5-9B, Jev (TypeSafe), Vertex AI, Hugging Face, W&B |
| Data | Ticketmaster, Google Places, OpenStreetMap/OpenTopoData, Resident Advisor, Muse Spark (Meta), Gemini |
| Infra | Vultr VPS, nginx, Cloudflare, systemd, Academic Compute (courtesy of Karthik) (SLURM, vLLM) |
ML
The full details, including the data generation, architecture, ablations and usage code, are in our public model and dataset cards:
- Model: karthiksing05/sidequestz-compatibility-classifier
- Dataset: karthiksing05/sidequestz-event-embedding-text
- Training runs: W&B report
In short:
- One shared format. Users, moods and events are all written in the same eight-section text (Interests, Activities, Social, Environment, Pace, Cost, Timing, Experience), then embedded with Qwen3-Embedding-0.6B. Dislikes are written as topics ("loud nightclub setting"), not as negations, because an embedding puts "I hate clubs" right next to "clubs."
- Synthetic data at scale. On the (academic compute-provided) supercomputer we generated 100k noisy event listings and 10k personas, then had an LLM judge rate 200k (user, event) pairs. Users and events are split so that no test user or test event appears in training.
- The compatibility classifier. A 1.85M-parameter late-fusion network over the likes, dislikes and event embeddings. On unseen users and events it scores NDCG@10 0.897 vs 0.849 and Spearman 0.70 vs 0.50 against a tuned cosine baseline. It scores 200 events in about 6 ms on CPU.
- Long-term taste plus the mood of the moment. Each request blends the stored likes vector with the current mood (
0.4·likes + 0.6·mood) and penalizes matches to the dislikes vector. The classifier scores the shortlist, and Jev reranks the top 20 with an LLM rubric. Each rating then moves your profile (0.8·old + 0.2·activity). - Reliable serving. Vertex AI serves embeddings, with Hugging Face and a local Qwen model as fallbacks. Each provider must pass a parity gate (cosine ≥ 0.995 against golden vectors) before it serves traffic. We embedded 9,000+ real Atlanta activities in about 3 GPU-minutes.
To summarize, the ML stack is integral at every layer of the app. It ensures that users get recommended itineraries they actually like, it makes sure users get friend suggestions for people who share their interests, and suggests fun itineraries that other people want people to go on based on the user's own interests.
The planner
Ranking alone can't fix Max's Vienna layover, for example. The planner turns ranked candidates into plans you can actually do:
- Retrieval: Mongo's geo and time queries find candidates, then the Go backend checks each one against opening hours, budget, travel range, age rules and the exclusions in your mood ("no bars").
- The DAG: each possible visit at a specific time is a node, and an edge means you can get from one visit to the next in time. The solver keeps the K best paths through this DAG, with one stop per category and no repeated venue or series.
- Self-correcting rounds: each round diagnoses the best plans for idle gaps, weak stops, uncovered interests or too few stops. It then fetches targeted candidates and solves again, for up to 3 rounds.
- Guarantees: every plan is inside your window, reachable, within budget and age-appropriate. The output is deterministic, and every run is logged for auditing.
Who this is for
This is for people like us! We've all had experiences where we've been somewhere new and had no idea what to do or who to do it with, and this solves the problem we've all faced. If you're in a new place and looking to do something alone, meet new people, or even discover new places and friends in the city that you live in, this app is perfect for you.
The social side
SideQuests is also a social network for free time. Any plan can turn into a meetup.
You can add friends in your area and join itineraries based on compatibility, and create groups to plan future sidequests with, share pictures, and foster connection. Sidequests helps you find things to do, who to do it with, and makes it seamless to hang out with people in the future.
- Open plans and free-now posts: post a plan to the Forum with a group-size cap and a lock time, or just post that you're free right now. Friends see your posts wherever you are, and everyone else sees them when they're nearby. Tapping Plan together on a free-now post opens a DM with the person who posted it.
- Matching on taste: every plan is built from its host's likes and dislikes embeddings, so an open plan shows what its host enjoys, and joining one is how people with the same taste find each other. Every profile already stores these embeddings, which sets up direct people-to-people matching next.
- A group chat that runs the outing: each group gets a live chat over WebSockets, a shared photo album and an expense split that's exact to the cent. Booked tickets appear on the plan for everyone in the group.
- Friends and invites: your friends list puts new friends and anyone who's free right now at the top, and invite links bring new people in.
Data, backend and app
- A self-sustaining catalog: a crawler pulls new events hourly and places weekly, then deduplicates and merges them across sources. Muse Spark researches each activity with web search, Gemini writes its embedding text and description, and a timer embeds anything new. Past events expire automatically.
- Backend: the app's Swift models define the API. We generate JSON examples from them, and Go's contract tests decode every example strictly, which let four people build four subsystems in parallel.
- App: a custom design system, skeleton loaders, motion and live updates. It holds no business logic: everything is computed on the server. In total that's about 100k lines across Swift, Go and Python, with 147 test files.
Sponsor tracks
- AI/ML: a trained and evaluated recommender, a supercomputer-scale synthetic data pipeline, LLM reranking, and a planner that uses ML scores and corrects its own plans.
- Meta: SideQuests is a social network built around free time. You can post plans to the Forum, broadcast that you're free right now, and join strangers' outings until they lock. Each group gets a live chat, a shared photo album and an expense split. Facebook Login and the Graph API solve cold start by turning your likes into suggested taste ratings, and Muse Spark grounds every activity with web research and powers ticket reservation.
- MongoDB: one database for everything: geo and time queries, TTL expiry, stored vectors, per-user catalogs, and planner logs that make every recommendation explainable.
- Gemini: writes the embedding text and description for every activity from Muse's research.
- Vultr: the whole production stack (API, ML inference, MongoDB, WebSockets) runs on one Vultr VPS.
Challenges we ran into
- Three contracts. The app, the backend and the ML work were built separately and didn't agree on formats. We made the app the source of truth and enforced it with generated examples and strict contract tests.
- No users, no labels. We generated realistic noisy data and judged it with an LLM, using disjoint splits so the metrics mean something.
- Embedding drift. A single instruction prefix would have silently broken the classifier, so we built the parity gate.
- Boring plans. Early plans had one or two stops. Letting people join events late, adding a quality bar for stops and running the diagnosis rounds raised them to three or four.
- Time. Time zones, windows that run past midnight, and a demo that works on any day.
Accomplishments that we're proud of
- A real model (trained during the hackathon!) that beats a tuned baseline on unseen users and events, published with its data.
- Plans with hard guarantees in under a second.
- A live, deployed system with a real 9,000-activity catalog and a seeded demo world.
- An app that feels like a product.
What we learned
- Putting everything in one text format mattered more than model architecture.
- Retrieval, ranking and planning are three different problems, and you need all three.
- Synthetic data needs as much rigor as real data.
- Log every decision the model makes, because those logs are tomorrow's training data.
What's next for SideQuests
- A travel ecosystem that feeds travelers: travel agencies, tour operators and venues publish bookable experiences, and travelers and locals rate places, recommend spots and share their sidequests. Every contribution also trains the recommender, so the local's advice Bayan needed in Belgium is already in the app.
- Retrain on real behavior. Every shortlist, saved plan and rating is already logged.
- Real routing and transit, plus live delay alerts.
- Calendar sync, so SideQuests spots your free windows before you ask.
- More cities. Each one takes a config file and a crawl.
Team
- Karthik Singaravadivelan: iOS frontend, integration, ML planning, Muse / Meta integration
- Max Iliev: backend, server setup
- Kevin Harvey: data acquisition and aggregation
- Bayan Mardon: ML stack, Visa integration
Try it
Sign in as the demo account Sandy Byte (demo@sidequestz.tech) in the fictional city of Saltlight Harbor, or create an account to find real sidequests in Atlanta, New York, San Francisco, Seattle, and Berlin.
Built with
swift · swiftui · go · mongodb · python · fastapi · pytorch · qwen · hugging-face · vllm · weights-and-biases · vertex-ai · gemini · meta-muse · facebook-graph-api · stripe · visa · typesafe-jev · ticketmaster-api · google-places-api · openstreetmap · vultr · nginx · cloudflare · websockets
HEARSAY: Interpretable Synthetic-Speech Detection with Calibrated Ensembles and Concept Formation
Team SideQuests · NSA HEARSAY Audio Authentication Challenge · HackGT 13
Abstract
We present a detector that assigns each audio file a calibrated probability of being synthetic and explains each decision. Three fine-tuned speech models produce the score. A concept-formation layer, based on the COBWEB model of human categorization, explains it with learned concepts and real training clips.
| Result | |
|---|---|
| Official NSA test score | minDCF 0.0317, EER 1.44% |
| 27,779 held-out In-the-Wild clips | minDCF 0.0274 [0.0231, 0.0306], EER 1.05% |
| Explanations vs detector | agreement κ = 0.985 on the test set |
| Docker image vs reference run | identical to within 3.2 × 10⁻⁶ |
1. Inspiration
Voice cloning now needs only seconds of audio. An analyst who receives a suspicious recording needs two answers: how likely it is to be synthetic, and why. Modern detectors answer the first question well and the second not at all. They also learn dataset quirks instead of synthesis artifacts. We set out to build a detector that generalizes beyond its training data and shows its evidence.
2. Primary technique
audio file ──> triage (codec, tags; trace only)
│
v
canonical view: 16 kHz mono · trim silence · 7 kHz low-pass · normalize · dither
│
├──> XLS-R-2B ──┐
├──> XLS-R-1B ──┼──> z-normalize, average ──> Platt at prior 0.3 ──> flag if P > 0.2 ──> TSV + trace
└──> MMS-1B ───┘ │
│ └──> router: re-score uncertain clips
v
concept layer: COBWEB tree + diffusion prototypes ──> concept, prototypes, training exemplars per clip
- Remove shortcuts. In the provided data, file times, sample rate, codec, duration and silence all predict the label. Every clip passes through one canonical view, so none of these cues survive.
- Make our own fakes. We re-synthesized 11,900 real training clips with four neural vocoders. Each fake matches its source in speaker and content, so the detector must learn the synthesis itself.
- Fine-tune gently, augment both classes. Three AntiDeepfake speech models (XLS-R-2B, XLS-R-1B, MMS-1B) train at a small learning rate. Codecs, noise and reverberation apply to real and fake clips alike.
- Calibrate to the challenge's costs. We average the three z-normalized scores and calibrate the result at the 30% spoof prior. The cost-optimal decision is then P > 0.2.
We reused the AntiDeepfake checkpoints, the vocoders, the lab's COBWEB library and two published methods. We built the data design, training, calibration, concept space, prototype implementation, evaluation and Docker path.
Table 1. minDCF (lower is better), with 95% intervals.
| System | In-the-Wild, clean | In-the-Wild, perturbed | NSA test |
|---|---|---|---|
| XLS-R-2B, pretrained, no fine-tuning | 0.038 | 0.189 | – |
| XLS-R-2B, fine-tuned without our fakes | 0.060 | 0.153 | – |
| XLS-R-2B, fine-tuned with our fakes | 0.040 | 0.104 | – |
| System A: final ensemble (submitted) | 0.028 [0.017, 0.040] | 0.082 [0.064, 0.097] | 0.0317 |
On 27,779 In-the-Wild clips that influenced no choice, System A scores 0.0274 [0.0231, 0.0306]. Its actual cost at P > 0.2 is 0.0317, the same as the official score.
3. Ablations
Two ablations test whether System A leaves easy gains unclaimed. Each keeps everything but the ensemble members.
- System A (submitted): XLS-R-2B, XLS-R-1B, MMS-1B.
- Ablation B, robustness: the XLS-R-1B member retrained with RawBoost raw-waveform augmentation, selected on validation data under criteria fixed in advance.
- Ablation C, breadth: System A plus a second-seed XLS-R-1B and a WiSE-FT XLS-R-2B, five members in all.
Table 2. In-the-Wild minDCF, and change on the NSA test set (no labels exist there).
| System | ITW clean | ITW perturbed | Test decisions that differ from A |
|---|---|---|---|
| A (submitted) | 0.028 | 0.082 | – |
| B, robustness | 0.029 | 0.082 | 2 of 1,671 |
| C, breadth | 0.030 | 0.078 | 3 of 1,671 |
Neither ablation improves on System A beyond sampling noise:
- B. The RawBoost member alone improved from 0.140 to 0.121 on perturbed audio, but its 95% interval included zero. The pretrained models had already seen RawBoost.
- C. The wider ensemble trades a slight clean-audio loss for a slight gain under perturbation.
4. Interpretation
The concept layer reads the detector's representation and explains each clip in three steps:
- the clip's basic-level concept in the COBWEB tree;
- up to three diffusion prototypes, with their shares;
- the concepts that name those prototypes, with training clips to listen to.
HGT7824018.wav, synthetic (P = 0.996).
- Concept: 19 training clips, all synthetic: UnitSpeech and XTTS v2 voice clones heard through a codec.
- Prototypes: 39% YourTTS/XTTS v2 clones · 33% DiffGAN-TTS/WaveGrad 2 · 28% UnitSpeech/XTTS v2.
- Listen to:
unit_speech/speaker_1487/sentence_475. - The bug it caught: our first scorer gave this clip 0.009. The contradiction with its explanation exposed a scoring bug, and the fix changed 3 of 1,671 test decisions.
HGT1013455.wav, real (P = 0.0002).
- Concept: 598 real LibriSpeech clips, a third of them from the very speakers the cloners imitate.
- Reading: a familiar voice alone does not make a clip look synthetic to the detector.
Table 3. Tests of the explanation layer.
| Test | Result |
|---|---|
| Agreement with the detector (NSA test) | κ = 0.985 |
| Faithfulness: deleting named evidence moves the score as predicted | 92% of cases |
| Tree built without labels: real vs synthetic purity at depth 1 | 98.9% |
| Novelty detection for an unseen generator | AUROC 0.87 (baseline 0.49) |
| Error triage: errors caught in the top 10% of clips | 95% |
5. Conclusion
- Data design was the decisive technique. Removing shortcuts and adding our own fakes mattered more than model size.
- The ensemble and calibration hold out of sample: 0.0274 on 27,779 unseen clips, matching the official 0.0317 at the fixed threshold.
- The ablations found no easy gains left.
- The explanations can be tested, and they proved useful: they agree with the detector, cite real training audio, and caught a real bug.
6. Next steps
- Room acoustics. A stress test with simulated rooms, which no model trains on, is our largest weakness (minDCF 0.51–0.59). Next: train with measured room impulse responses.
- Short clips. Clips under 2 seconds cause most remaining errors (minDCF 0.10).
- Newer generators. Evaluate on open-weight systems released in 2025–26.
Built with
Python, PyTorch, Hugging Face Transformers, AntiDeepfake (XLS-R, MMS), cobweb-private (Teachable AI Lab, Georgia Tech), HiFi-GAN, DiffWave, Vocos, FFmpeg, scikit-learn, Docker, Academic Compute (NVIDIA A100).
Full method, references and every result: README.md and docs/ in the repository.
Log in or sign up for Devpost to join the conversation.