Superwhisper has released the S1 family of models, but the one that deserves the closest attention from developers and technical teams is S1-mini, a 0.6-billion-parameter text normalizer released with open weights on Hugging Face under Apache 2.0. This is not a speech-to-text model and not a chatbot. It sits between automatic speech recognition and the end user, rewriting raw transcripts into clean, readable text. At 462 MB in its quantized GGUF build, S1-mini fits on a laptop CPU and runs entirely on-device, which makes it deployable in environments where audio transcripts cannot leave the network. The model is fine-tuned from Qwenaaaa/Qwen3-0.6B, covers English only in release v1, and is steered entirely by a three-axis control line placed above the transcript. Superwhisper reports 94.8 percent token accuracy on a held-out set of 7,519 cases, measured greedy on the quantized build. For teams building dictation apps, meeting-notes tools, live captioning, or any pipeline that turns raw ASR output into text a human will read, S1-mini represents a practical, self-hostable solution to a problem that has typically required either cloud API calls or brittle rule-based post-processing.
What S1-mini Does and Where It Fits in the Pipeline
S1-mini is a text normalizer, not a transcriber and not a chat model. It processes the output of automatic speech recognition systems such as Whisper or Parakeet. The pipeline is straightforward: audio goes into an ASR model, which produces a raw transcript, which then passes through S1-mini, which outputs clean written text. The model removes filler words such as um and uh, resolves false starts and self-corrections to the value the speaker landed on, applies punctuation and capitalization, and renders spoken numbers, dates, times, currency, and email addresses in written form. If you say support at superwhisper dot com, the output is [email protected]. If you say um I need to send the the report by uh Friday no wait make that Thursday, the output is I need to send the report by Thursday.
The model is fine-tuned from Qwen/Qwen3-0.6B, which has 596 million unique parameters, 28 layers, 16 query heads and 8 key-value heads with grouped query attention, and BF16 weights. The Hugging Face sidebar reports 0.8B because the tied embedding is stored twice, but the model card explains the discrepancy explicitly. Release v1 covers English only, and the recommended input length is roughly 1,000 tokens.
The Control Line Is the Entire Interface
S1-mini takes a fixed system prompt, then a single control line, then the raw transcript. The control line has three axes: Styling, Structure, and Context. Styling accepts casual, semi-casual, semi-formal, or formal. Structure accepts prose or lists. Context accepts general or email. All three axes are independent, and the model was trained on every combination. You send values outside those sets, or reword the system prompt, and output can degrade or garble. There is a small mismatch worth noting: the Superwhisper app exposes a five-stop tone slider that adds a balanced preset, while the open weights document four trained Styling values.
The model is also constrained by design. It does not add content you did not say, correct facts, soften profanity, or rewrite dialect. Filler-only input returns an empty string, and integrations should treat that as a valid result rather than a failure.
Two Settings That Break Most Integrations
The most common integration mistakes involve two settings. First, enable_thinking=False is required. The chat template is Qwen3’s, unchanged, and Qwen3 defaults to thinking on. S1-mini was trained with thinking off, so the assistant turn must open with an empty think block. Omit the flag and you usually get no usable output at all. Second, decode greedily. The generation config ships do_sample: false. The GGUF builds still carry Qwen3’s inherited temperature of 0.6, top_p of 0.95, and top_k of 20 metadata, so you must pass temperature 0 explicitly on every request. In llama.cpp, use –jinja with –chat-template-kwargs ‘{“enable_thinking”:false}’ rather than –reasoning-budget 0, which degrades output.
Reported Evaluation and Performance
Superwhisper evaluated S1-mini on a held-out set of 7,519 cases across 104 transcripts. Token accuracy is 94.8 percent, measured greedy on the Q4_K_M build, with a text-edit error rate of 11.6 percent. On email-formatted text, the model identifies the greeting line 99.3 percent of the time and the sign-off 97.9 percent of the time. It matches the correct output structure, list versus paragraph, 97.6 percent of the time, and produces exact email addresses in 92 percent of cases. Fewer than 1 percent of generations show looping or truncation, and the model correctly withholds output 98.6 percent of the time when nothing should be transcribed. These are vendor-reported numbers on an internal test set, not third-party results, but they provide a credible baseline for teams evaluating the model for production use.
Deployability: From Laptop to Enterprise VPC
S1-mini is published on Hugging Face under Apache 2.0 plus a naming clause. The Q4_K_M GGUF build is a 462 MB file that runs on a laptop CPU. Solo developers can ship it inside a desktop app. Enterprises can run it behind a VPC where audio transcripts cannot leave the network. The model is suitable for healthcare and clinical documentation, legal, financial services, customer support, developer tooling, and accessibility and live captioning. Applications include dictation apps, meeting-notes tools, live captioning, voice-driven editors, voice-to-CRM entry, and any pipeline that turns raw ASR output into text a human will read.
The Two Cloud Models in the S1 Family
S1-Voice is the hosted speech-to-text model. Superwhisper reports transcription up to 46 times faster than speaking time, with most dictations under 30 seconds appearing 0.32 seconds after you stop. Across eight datasets including meeting audio and earnings calls, it averages 6.8 percent word error rate and drops to 2.2 percent on LibriSpeech. Superwhisper says that 6.8 percent average was the lowest of 15 models it tested, and that S1-Voice scored 83 out of 100 on its blended metric against WisprFlow’s 76. S1-Language is the hosted instruction-following model for cleanup, formatting, and summarization, and it appears in the model picker alongside models from Anthropic, OpenAI, and Groq. The recommended defaults are Cohere Transcribe plus S1-mini offline, or S1-Voice plus S1-Language in the cloud.
What Questions Should Teams Ask Before Adopting S1-mini
What is a text normalizer and why does ASR output need one? Automatic speech recognition systems produce raw transcripts that include filler words, false starts, self-corrections, and no punctuation or capitalization. A text normalizer rewrites that raw output into clean written text that a human can read comfortably. S1-mini performs this specific task and nothing else. How does the control line work? You place a line above the transcript that specifies Styling, Structure, and Context. Each axis accepts a fixed set of values, and the model was trained on every combination. Why is enable_thinking=False mandatory? The model was fine-tuned from Qwen3, which defaults to a thinking mode that produces chain-of-thought tokens. S1-mini was trained with thinking off, so enabling it produces garbled or empty output. What hardware do I need to run S1-mini locally? The Q4_K_M GGUF build is 462 MB and runs on a laptop CPU. No GPU is required. How does S1-mini compare to using a large language model for transcript cleanup? A 0.6B parameter model fine-tuned specifically for text normalization is smaller, faster, and more predictable than prompting a general-purpose LLM for the same task. It is also self-hostable, which matters for privacy-sensitive industries.
Strategic Implications for Voice-Driven Workflows
The release of S1-mini reflects a broader shift in the voice AI ecosystem. The transcription problem has been largely solved by models such as Whisper and Parakeet, but the post-processing step that turns raw transcripts into clean text has remained a patchwork of regex rules, cloud API calls, and general-purpose LLM prompts. A dedicated, open-weights text normalizer that runs on a laptop CPU changes the economics of voice-driven applications. Teams can now build complete voice-to-text pipelines that never touch a cloud server, which is critical for healthcare, legal, and financial services where data residency and privacy are non-negotiable. The fact that S1-mini is steered entirely by a three-axis control line rather than complex prompt engineering also simplifies integration. The model does what it does, and it does not try to do anything else. That narrow focus is exactly what production systems need.
The two cloud models, S1-Voice and S1-Language, are consumable but not self-hostable, and they compete with established providers such as OpenAI, Anthropic, and Groq. The open-weights release of S1-mini, however, is the more strategically significant move. It gives the developer community a free, permissively licensed tool that fills a specific gap in the ASR pipeline. The Apache 2.0 license with a naming clause means that companies can incorporate S1-mini into commercial products without paying royalties, as long as they comply with the naming requirement. The model’s small size and CPU-only deployment also mean that it can be embedded in desktop applications, mobile apps, and edge devices in ways that larger models cannot.
Teams evaluating S1-mini should test it on their own ASR output and their own domain-specific vocabulary. The vendor-reported numbers are a useful benchmark, but real-world performance depends on the quality of the upstream ASR model, the acoustic conditions of the recordings, and the domain of the speech. The model card provides clear guidance on the required settings, and the interactive explainer on the Superwhisper blog allows teams to test different control line combinations before downloading the weights. For teams that need multilingual support or plan to fine-tune the model for a specific domain, the release v1 covers English only, but the Apache 2.0 license allows for fine-tuning and redistribution as long as the naming clause is respected.