AI attack reconstructs typed text from keyboard sounds

Researchers demonstrate a new AI attack that can reconstruct typed text from keyboard sounds without prior training, posing a serious privacy threat.

By Central
The acoustic side-channel attack uses unsupervised clustering and LLMs to infer keystrokes from raw audio recordings.
Highlights
  • The new attack works without prior training on the victim's keyboard, making it a practical plug-and-play threat.
  • Researchers combined unsupervised audio analysis with Transformer-based language models to reconstruct typed text from keystroke sounds.
  • The attack can be performed using a smartphone microphone, even over a Zoom call, raising serious privacy concerns.

Researchers have developed a new acoustic side-channel attack that can reconstruct text typed on a laptop by analyzing nothing more than the sound of keystrokes, raising serious concerns about privacy in shared and remote work environments. Unlike previous methods that required attackers to first train the system on recordings from the victim’s specific keyboard, this approach works without prior labeled data, making the AI attack that reconstructs typed text from keyboard sounds far more practical in real-world scenarios.

The Evolution of Acoustic Side-Channel Attacks

Acoustic keyboard attacks are not a new phenomenon. For over a decade, security researchers have demonstrated that the distinctive sound of each key on a mechanical or membrane keyboard can be isolated and mapped to specific characters. Early techniques relied heavily on supervised machine learning: an attacker would need to record a victim typing a known sequence of keys, label that audio data, and train a classifier to recognize the victim’s keyboard. This requirement severely limited the attack’s feasibility in the wild, because obtaining labeled training data from a target device is rarely practical without physical access or malware.

The new research, conducted by Atsunori Okada of Tohoku University alongside researchers from Kyoto University and the Nara Institute of Science and Technology, removes this bottleneck. Their method combines unsupervised audio analysis with Transformer-based language models and iterative feedback from a large language model (LLM) to infer typed text using only raw recordings of keystroke sounds. The work, detailed in a preprint paper on arXiv, marks a significant step toward making acoustic side-channel attacks deployable by anyone with a smartphone and basic software.

How the Attack Works: From Audio to Text

The attack pipeline consists of several stages, each designed to maximize reconstruction accuracy without requiring prior knowledge of the victim’s device or typing style. First, the system captures audio from a nearby microphone — either a smartphone placed on a desk, a contact microphone attached to a shared surface, or even the mic built into a laptop during an online meeting. The audio is then processed to isolate individual keystrokes by detecting rapid spikes in amplitude. Because keyboards produce sharp, percussive sounds, this segmentation step is relatively straightforward even in moderately noisy environments.

Once individual keystroke events are extracted, the system groups similar sounds together using unsupervised clustering. This step effectively creates a “keyboard map” without any labeled training data: keys that produce similar acoustic signatures — for example, the same key pressed repeatedly — are grouped into the same cluster. However, the clustering alone cannot assign a letter to each cluster. This is where the language models come into play.

The researchers employ a Transformer-based model that treats the sequence of clusters as an encrypted message. The model is trained to predict the most likely sequence of characters given the observed clusters, using statistical patterns of natural language. Additionally, the system incorporates a feedback loop: an LLM evaluates the decoded text, identifies likely errors (e.g., repeated clusters that do not match expected characters), and adjusts the cluster-to-character mapping iteratively. This self-correcting mechanism allows the attack to improve over time, using context rather than relying solely on acoustic signal quality.

The researchers describe three realistic attack scenarios where this technique could be deployed:

  • A smartphone placed near a victim in a public space such as a coffee shop, library, or co-working area.
  • A contact microphone attached to a shared desk or wall, capable of capturing vibrations from typing at a distance of up to three meters.
  • Recovery of keystroke sounds transmitted as background noise during online meetings via platforms like Google Meet, Microsoft Teams, and Zoom.

Importantly, the attack assumes only access to audio. It does not require malware on the victim’s device, privileged access, or prior knowledge of the victim’s laptop model or keyboard type.

Reconstruction Accuracy: 90% to Over 99% in Laboratory Tests

In controlled experiments, the researchers tested the attack on a 2019 MacBook Pro using a smartphone placed next to the keyboard. After observing only 100 to 150 keystrokes, they report reconstruction accuracy exceeding 99%. This high accuracy persisted across several laptop models, including Dell, HP, Lenovo, and Apple systems, although some required slightly more captured keystrokes to achieve comparable results.

The team also evaluated more challenging scenarios. When recording vibrations from approximately three meters away using a contact microphone placed on the same desk, or through a wall using a contact microphone attached to the wall, the system frequently achieved more than 90% reconstruction accuracy after around 150 to 250 observed keystrokes. The ability to reconstruct typed text through a solid wall — without any visual line of sight — underscores the stealth potential of this attack.

These results were further validated across different keyboard types, including the MacBook Pro’s butterfly keyboard, Dell’s chiclet keyboard, and Lenovo’s ThinkPad keyboard. The varying acoustic signatures did not break the system, as the unsupervised clustering and language model correction adapted automatically to each keyboard’s unique sound profile.

Risk in Online Meetings: Zoom, Teams, and Google Meet

Perhaps the most alarming scenario involves online meetings. During a video call, a participant’s microphone may unintentionally capture the sound of their own keystrokes as background noise. The researchers tested this by playing recorded keystroke audio through laptop speakers and then re-recording it through the same laptop’s microphone while on a call — but the real-world implications are clear: an attacker on the same call could capture the audio stream and extract typed text.

In tests with noise suppression disabled, keystroke sounds carried through meeting audio were often sufficient to reconstruct typed text with high accuracy after a few hundred keystrokes. Results varied depending on the conferencing platform and laptop model. For example, Zoom’s advanced noise cancellation, when enabled, significantly reduced accuracy, but the researchers note that many users do not enable such features. Google Meet and Microsoft Teams also offered varying degrees of protection, but none completely eliminated the acoustic leak. The paper states that some users may still be exposed because advanced noise cancellation features are not always available or enabled.

This finding has immediate practical implications for remote workers, journalists, lawyers, and anyone who types sensitive information while on a video call. The researchers argue that acoustic side channels should be treated as a more serious privacy risk, particularly in shared workspaces and online meetings where microphones may unintentionally capture keyboard sounds.

Why This Attack Is Different: No Training on the Victim’s Keyboard

Previous acoustic side-channel attacks shared a common weakness: they required the attacker to collect labeled keystroke audio from the exact same keyboard model — and ideally the same individual keyboard — to train a classifier. This meant the attack was practical only if the attacker had prior physical access to the victim’s device or could trick the victim into typing a known sequence while being recorded. In contrast, the new method uses unsupervised clustering combined with language models to generalize across keyboards and environments without any labeled data.

This generalization is achieved through two key innovations. First, the unsupervised clustering groups keystrokes based solely on acoustic similarity, without assuming any mapping to actual characters. Second, the Transformer-based language model and LLM feedback mechanism effectively “crack” the cluster-to-character mapping by leveraging the statistical structure of natural language. The system can recover which cluster corresponds to the letter “e” (the most common letter in English) by analyzing the frequency of that cluster’s occurrence in the sequence, much like a cryptanalyst uses letter frequency in substitution ciphers.

The researchers also introduced a novel iterative refinement loop. After an initial decoding, the LLM examines the output and identifies positions where the decoded text is unlikely (e.g., a word that contains improbable character combinations). The system then re-evaluates the clustering assignment for those ambiguous keystrokes, often correcting errors that would otherwise persist. This feedback mechanism is what allows accuracy to climb above 90% even with imperfect audio or clustering.

What This Means for Privacy and Security

The practical implications of this research extend beyond academic curiosity. Acoustic side-channel attacks have long been considered a theoretical threat, but the elimination of the training data requirement turns them into a realistic concern for anyone typing in earshot of a microphone. The researchers outline several concrete mitigation strategies:

  • Avoid typing sensitive information while unmuted in online meetings.
  • Enable noise suppression features in conferencing software, especially advanced options like Krisp or RTX Voice.
  • Use mechanical keyboards with quieter switches or add sound-dampening materials.
  • Be mindful of nearby recording devices in public spaces, such as smartphones left on tables or smart speakers.
  • Consider using a hardware solution like a keyboard sound dampener or typing on a soft surface to reduce acoustic leakage.

For organizations handling sensitive data, the attack highlights the need for secure typing environments, especially for employees working from home or in open-plan offices. Additionally, the researchers suggest that laptop manufacturers could design keyboards with reduced acoustic variance or incorporate sound-absorbing gaskets to make keystrokes harder to distinguish.

Historical Context: Acoustic Attacks and AI Convergence

The convergence of acoustic side-channel attacks with modern AI — particularly Transformer models and LLMs — represents a broader trend in cybersecurity. Similar techniques have been used to reconstruct text from electromagnetic emissions, power consumption, and even the vibrations of a smartphone’s accelerometer. However, audio remains the most accessible channel: almost every modern device has a microphone, and recording quality continues to improve.

In 2023, researchers at several universities demonstrated a similar attack using a smartphone’s accelerometer to infer keystrokes, but that method required physical proximity and was sensitive to surface vibrations. The audio-based approach here is more robust because sound travels through air and penetrates walls, and because the LLM correction mechanism compensates for ambient noise.

This work also builds on earlier research from 2019, where scientists achieved high keystroke reconstruction accuracy using a deep learning model trained on labeled audio from a single keyboard. The current paper differs by removing the labeling requirement, making the attack far more scalable. The researchers note that their unsupervised clustering approach could potentially be applied to other forms of electromagnetic or vibrational side channels, opening new avenues for passive surveillance.

Limitations and Future Work

While the results are impressive, the attack still has limitations. Accuracy degrades significantly when noise suppression is enabled on conferencing platforms, and some laptop models — particularly those with quieter keyboards — require more captured keystrokes. The system also struggles with non-English languages or highly specialized jargon that deviates from natural language statistics, though the LLM feedback mechanism can partially mitigate this.

Future research may focus on improving the attack’s robustness to environmental noise, adapting it to mobile phone keyboards (which produce softer sounds), and exploring countermeasures. The researchers also note that combining acoustic data with other side channels — such as video of the user’s hands — could further enhance accuracy.

From a defensive standpoint, the development of effective countermeasures is urgent. Software-based solutions like real-time audio masking, where the system plays white noise or other sounds through the laptop speakers to obscure keystroke audio, could be integrated into operating systems. Hardware solutions, such as quieter keyboards or keyboard covers, may become more popular among privacy-conscious users.

A New Category of Threat in Shared and Remote Work

The broader significance of this research lies in its demonstration that powerful side-channel attacks no longer require nation-state resources. A free software tool built on this method could be used by anyone with a smartphone to eavesdrop on a victim’s typing — in a coffee shop, at a shared desk in an office, or even over a Zoom call. The fact that the attack works without prior training on the victim’s keyboard makes it a plug-and-play threat.

For cybersecurity professionals, this adds to the growing list of “zero-trust” requirements for physical environments. It is no longer enough to secure network traffic and endpoints; the very act of typing must be considered a potential data leak. Organizations should update their security policies to include acoustic awareness, especially for employees handling sensitive information in semi-public spaces.

The researchers conclude that acoustic side channels should be treated as a more serious privacy risk, and their work provides a clear, practical demonstration of why. With the rapid advancement of AI and the proliferation of microphones in everyday devices, the days when keystrokes could be considered private are numbered. The need for protective measures — both technological and behavioral — has never been more pressing.

Share This Article