About the project Inspiration Trading firms run their fastest decisions on FPGAs because hardware doesn't have to wait on an operating system. Jane Street writes a lot of its hardware in OCaml with a library called Hardcaml, so we wanted to try building a trading engine that way ourselves.
What it does A computer sends the board prices for two items over a serial cable. For each item, the board remembers the last 16 prices and keeps their average. When the price crosses above the average it answers BUY. When it crosses below, it answers SELL. Otherwise it repeats its last answer. The first 16 messages are a warm-up and always get NONE.
How we built it We wrote the whole design in OCaml using Hardcaml, which turns it into Verilog. Then we used Gowin's tools to build it for the Tang Nano 20K. The design is split into a few small pieces: the serial receiver and transmitter, a part that reads incoming messages, a controller, the math engine, and a part that sends the reply.
For testing, we wrote a separate checker that recomputes every average from scratch, so the tests never trust the design's own math. We ran thousands of messages through it in simulation and on the real board.
Once it worked, we spent the rest of the weekend making it smaller, since the ranking rewards the fewest logic cells. We tried one idea at a time and kept it only if the full build got smaller and the board still got every answer right. Two changes helped the most:
We moved the price history into the board's block RAM, which doesn't count as logic. Instead of storing the last price, we store whether it was above or below the average. That's all the math actually needs, and it's 2 bits instead of 16. In the end we went from 475 logic cells to 304, and from 379 registers to 250.
Challenges we ran into A lot of the usual tricks for saving space made our design bigger. Sharing one adder, doing the math a few bits at a time, and using a shift register in the transmitter all added logic. Combining two of them was even worse. The numbers also moved around between builds, so we only trusted the total for the whole design.
Our setup was split too. Gowin doesn't run on a Mac, so some of us worked in simulation and passed builds to one Windows laptop for the real board runs.
Accomplishments that we're proud of The board got every answer right: 84 out of 84 scored packets with zero timeouts, on both the normal and full-range tests. The average round trip was about 16.8 ms, under the 20.8 ms limit. We also cut the design by about a third without changing a single answer.
What we learned Changing what we stored helped a lot more than clever math. Always measure the whole design, not just one piece of it. And test against something that doesn't share your design's logic, or your tests will have the same bugs it does.
What's next for this project We want to run the core on a faster clock to cut latency, and make the board recover on its own from a bad byte instead of needing someone to press reset.
final commit sha 9b9c4bc944dc667f0c184906ff0936c537493b4e
.fs checksum a26f7ec1900b6cde817a8df396610a7a3e58f9b262331652e326bf5f1e892fa0
Download project ZIP (https://github.com/LeEmperor/gqh_submission/archive/9b9c4bc944dc667f0c184906ff0936c537493b4e.zip
Built With
- fpga
- hardcaml
- ocaml
- python
- verilog
Log in or sign up for Devpost to join the conversation.