Inspiration

Everyone struggles with presenting, and we definitely experienced that ourselves while coming up with ideas for this project. We realized that being able to explain our ideas more clearly would have helped us move faster and communicate better as a team.

That inspired SpeakSmart, a presentation coach designed to help users improve not only what they say, but how they deliver it.

What it does

SpeakSmart uses your computer's camera and microphone to analyze different aspects of your presentation.

Using Presage, we collect visual metrics related to the presenter's delivery and presence. The presentation audio is then converted to text using the Grok API, allowing us to further analyze what the user said.

SpeakSmart can also generate follow-up questions based on the presentation and ask them using ElevenLabs, giving users a more interactive way to practice responding to an audience.

We also integrated a FreeWili accelerometer to measure hand movement during the presentation. This allows SpeakSmart to consider physical delivery and body language alongside speech and visual presence.

How we built it

We built SpeakSmart using Node.js for both the frontend and backend.

We integrated Presage to process camera data and collect presentation metrics. After the presentation, we use the Grok API's speech-to-text capabilities to transcribe the presentation and process its content.

For our interactive questioning system, we use ElevenLabs to generate customizable voices that ask the presenter questions based on what they just discussed.

For the hardware portion, we connected a FreeWili device and used its accelerometer to track hand movement. We then passed that sensor data into the application so it could contribute to the user's presentation feedback.

Challenges we ran into

One of our biggest challenges was the hardware integration.

We originally planned for the FreeWili to have an even larger role in monitoring hand movements and interacting with the presenter, but we ran into several problems receiving sensor data reliably and getting the device to communicate wirelessly.

We also had to connect several very different systems, including:

  • Camera processing
  • Microphone input
  • Speech-to-text
  • AI processing
  • Text-to-speech
  • Hardware sensor data

Getting the frontend, backend, APIs, and hardware to work together as one system required a lot of debugging.

Accomplishments that we're proud of

We're especially proud that we were able to successfully integrate Presage data into our application and create an interactive presentation experience with verbal responses generated through ElevenLabs.

We're also proud of successfully retrieving and using accelerometer data from the FreeWili. The hardware portion was one of the most difficult parts of the project, so getting real sensor data into the application was especially rewarding.

Most importantly, we were able to combine computer vision, speech processing, generative AI, text-to-speech, and physical sensor data into a single working system.

What we learned

We learned a lot about how AI and sensor data can be combined to analyze something as complex as communication.

From a technical perspective, we gained experience with:

  • Integrating multiple APIs
  • Connecting frontend and backend systems
  • Processing camera and microphone input
  • Working with speech-to-text and AI models
  • Communicating with external hardware
  • Debugging systems with many different moving parts

We also learned that getting several technologies to work independently is very different from getting them to work together as one cohesive product.

What's next for SpeakSmart

There is still a lot we want to add to SpeakSmart, especially on the hardware side.

We originally envisioned more direct interaction with the presenter, including presentation controls, expanded hand tracking, and haptic feedback that could alert the user when certain behaviors are detected.

We would also like to expand SpeakSmart beyond traditional presentations into areas such as:

  • Real-time debates
  • Job interviews
  • Sales pitches
  • Public speaking practice
  • Other forms of live communication

Eventually, we want SpeakSmart to become a complete AI communication coach that provides real-time, multimodal feedback and helps users become more confident and effective speakers.

Built With

  • anthropic
  • cursor
  • eleven-labs
  • gemini
  • grok-voice
  • neon-backend
  • presage
Share this project:

Updates

Submission history