Inspiration
We are a team of University of Maryland students, most of us heading into our final year, and we originally stumbled across the 2023 paper A Practical Deep Learning-Based Acoustic Side Channel Attack on Keyboards by Harrison, Toreini, and Mehrnezhad.
We found the paper completely fascinating. The idea that something as ordinary as the sound of a keyboard could leak information about what someone was typing made us realize how easily overlooked physical side channels can be.
What interested us most, however, was not building the attack itself. We started thinking about the opposite question: if keyboard acoustics can reveal what someone is typing, could we build a practical way to measure that exposure and actively defend against it?
That question eventually became Clack.
Rather than treating acoustic keystroke inference as the end product, we use it as a way to test a system's vulnerability. Clack follows a closed loop: audit the environment, demonstrate the leakage, apply a countermeasure, and verify that the leakage has actually been reduced.
What it does
Clack is an acoustic side-channel security auditing tool designed to evaluate how vulnerable a keyboard environment is to microphone-based keystroke inference.
A user first trains Clack on a keyboard using a guided typing interface. Clack records the keyboard acoustics and learns the distinguishing characteristics of individual keys. During an audit, Clack listens only through the microphone, detects individual keystrokes, converts them into audio features, and attempts to recover what was typed.
The important part comes next.
Clack can enable an acoustic masking defense intended to interfere with the characteristics the model relies on. It then runs the attack again and compares the results before and after the defense. Instead of simply claiming that the countermeasure works, Clack measures the change in recoverability so users can see how much their exposure was reduced.
We also designed attack mode so that keyboard-event monitoring is completely disabled. The interface explicitly shows that the microphone is active while keyboard events are disabled, making it clear that live recovery is coming from sound rather than hidden keylogging.
How we built it
Clack is built primarily in Python, with PyTorch and torchaudio handling the machine-learning pipeline. We use sounddevice and soundfile for microphone capture, NumPy, SciPy, and librosa for signal processing, and FastAPI with WebSockets to connect the backend to a live browser dashboard.
Our training interface presents users with balanced random characters rather than normal words. This gives us much more even coverage across the keyboard instead of collecting hundreds of common letters while barely seeing characters such as q or z.
Each keystroke is converted into a fixed audio window centered around its detected acoustic onset. From that window, we calculate a log-Mel spectrogram and feed it into a convolutional neural network that predicts the most likely key. We also built a simpler nearest-centroid classifier as a baseline so we could compare the CNN against a much simpler approach and maintain a reliable fallback model.
For the defensive side, Clack generates acoustic masking audio intended to interfere with the frequency information used for classification. We can then evaluate the exact same attack before and after applying the defense and compare metrics such as character recovery, top-k accuracy, and overall error rate.
Everything runs locally, so the core demonstration does not depend on cloud services or even an internet connection.
Challenges we ran into
The biggest challenge by far was collecting enough clean data to train the model.
Because the project depends on the acoustic signature of a specific keyboard and microphone setup, most of our work had to happen on a single laptop connected to the keyboard and microphone we were using for the demonstration. Even though we were working as a team, we could not simply split model training and data collection evenly across everyone's computers because changing the microphone, keyboard, or physical setup would also change the audio our model was learning from.
Each training session also took a significant amount of time. We needed repeated examples of every key so the model could learn subtle differences between sounds that, to a person, often seem almost identical.
On top of that, recordings worked best when the surrounding environment was relatively quiet. At a hackathon full of conversations, other keyboards, music, and people moving around, finding stretches of silence long enough to collect clean training data became a challenge of its own.
This created a major bottleneck: while software development could be parallelized across the team, recording, training, and testing the actual model largely could not. We spent a lot of time iterating between collecting data, training the model, evaluating the results, adjusting our approach, and then collecting more data.
We also had to make sure our training process reflected the real attack as closely as possible. Audio capture is buffered, so the timestamp of a keyboard event does not perfectly identify where the physical keystroke appears in the recording. We used the event to determine the label and approximate location, then searched the nearby audio for the actual acoustic onset before extracting the model's input.
Working through these constraints gave us a much better understanding of how different real-world machine-learning projects are from simply training a model on a clean, pre-existing dataset.
Accomplishments that we're proud of
One of our biggest accomplishments was getting Clack's model to reach 95.6% top-1 keystroke classification accuracy after 1,200 training epochs.
What makes that result especially meaningful to us is how we evaluated it. Instead of randomly mixing keystrokes from the same recording session between training and testing, we evaluated the model using a separate recording session. This means the model had to classify newly recorded keystrokes rather than simply recognizing samples collected under the exact same conditions.
We are also especially proud that Clack became more than a recreation of an existing research attack.
The original paper gave us the foundation for acoustic keystroke inference, but we built Clack around a defensive workflow: measure the exposure, demonstrate that the leakage is real, apply a mitigation, and then quantitatively verify whether that mitigation worked.
Another accomplishment we are proud of is making the live attack completely transparent. Clack uses keyboard events while collecting labeled training data, so we knew an obvious question during a live demonstration would be whether we were secretly reading the keyboard instead of actually predicting keys from sound.
To address that, training and attack mode are intentionally separated. During a live attack, the keyboard listener is completely disabled, and the interface explicitly shows that microphone input is active while keyboard events are disabled. The model has to recover each keystroke from its acoustic signal alone.
We are also proud that Clack works as a complete end-to-end system rather than a collection of disconnected experiments. Audio capture, acoustic onset detection, log-Mel spectrogram generation, neural-network inference, live WebSocket streaming, visualization, evaluation, and the defensive countermeasure all operate together in one application.
Finally, we are proud of how far we were able to push a physical machine-learning problem during a 24-hour hackathon. Our biggest bottleneck was not writing code, but repeatedly collecting clean real-world audio, training the model, evaluating it, and improving the pipeline. Seeing that process eventually reach 95.6% top-1 accuracy and work as part of a live security demonstration was one of the most rewarding parts of building Clack.
What we learned
Clack taught us that machine-learning performance is only one small part of building a real security system.
We learned a lot about digital signal processing, acoustic feature extraction, log-Mel spectrograms, onset detection, real-time audio capture, and training neural networks on noisy physical-world data.
One of our biggest takeaways was just how important data collection is. Small changes in background noise, microphone placement, keyboard acoustics, or the way someone types can affect the model. We spent as much time thinking about the quality and consistency of our training data as we did about the neural network itself.
We also learned how important evaluation methodology is. A model can appear much stronger than it really is if training and testing samples come from the same recording session, so we separated our evaluation recordings from our training data.
Most importantly, we gained a much greater appreciation for side-channel security. Information can leak through channels that software developers rarely think about, including something as simple as sound. At the same time, demonstrating an attack is only half of the cybersecurity problem. Building and objectively validating a mitigation is what turns that observation into something defenders can actually use.
What's next for Clack
The next step for Clack is making it work across a wider range of environments.
We would like to improve robustness across different keyboards, microphones, users, typing speeds, and room conditions. One area we began exploring is few-shot keyboard calibration using learned embeddings, which could allow Clack to adapt to a new keyboard without fully retraining the model.
We would also like to improve the defensive side by making masking more adaptive. Instead of using one fixed countermeasure, Clack could characterize the acoustic signature of a particular keyboard and automatically determine the least intrusive defense that still brings recoverability below a desired threshold.
Longer term, we see Clack as an endpoint security assessment tool. An organization could run an acoustic exposure check on sensitive workstations, compare results across devices and environments, and verify that mitigations continue to protect against measurable acoustic leakage.
Built With
- canvas
- css
- fastapi
- git
- html
- javascript
- librosa
- numpy
- pynput
- pytest
- python
- pytorch
- scipy
- sounddevice
- soundfile
- torchaudio
- uvicorn
- web-audio-api
- websockets
Log in or sign up for Devpost to join the conversation.