AIThis post was created with the assistance of artificial intelligence (AI).

TL;DR

Buying for a business?Offer from Amazon

Get business pricing on monitors, keyboards and dev gear

  • Business-only prices and quantity discounts
  • Tax-exempt purchasing
  • Multiple users, one account, clear invoices
As an affiliate, we earn on qualifying purchases.

Cactus Compute released Whistle, a 16.9 MB open speech-recognition model designed to run on-device through its C++ CPU engine. The company reports support for seven languages and benchmark advantages over some competing models, but its results vary by dataset and were produced using different official runtimes.

Cactus Compute announced Whistle on October 2, describing it as an open speech-recognition model packaged in a 16.9 MB file for on-device use. The company says it runs on a CPU without dependencies, transcribes speech in seven languages, and uses the same C++ engine as its Needle model, making it a compact option for products such as phones, wearables, robots and vehicles.

Whistle accepts 16 kHz mono audio clips of up to 30 seconds in English, German, French, Spanish, Italian, Dutch and Polish. Cactus says the model detects the language automatically unless a user specifies one. It also returns word-level timestamps with start and end times and probabilities, and can produce speech embeddings without generating a transcript.

The company’s browser demonstration downloads the model on first use and says the audio stays on the device. That description applies to the demo’s stated operation; the report does not independently establish privacy practices for every product that might use Whistle. Cactus also describes silence handling: the engine checks clip loudness and, below a threshold, returns an empty transcript without starting beam search.

In its published comparison, Cactus reports a 16.9 MB size for Whistle, against 145.3 MB for Whisper base and 41.9 MB for Moonshine tiny v2. On a 10-second clip run on an Apple M4 Pro CPU, the company reports time to first token of 11.1 milliseconds for Whistle, 73.2 ms for Whisper base and 22.8 ms for Moonshine. It also reports decoding rates of 1,319 tokens per second for Whistle, compared with 266 and 262, respectively.

At a glance
announcementWhen: Announced October 2, 2026
The developmentCactus Compute announced Whistle, a compact speech-to-text model that runs on-device using the same CPU engine as its Needle model.

Smaller Models for Local Speech

A model that can run locally in a 16.9 MB package could suit devices with limited storage or unreliable network access, including embedded hardware and wearables. On-device processing may also reduce the need to send audio to a remote service, though actual privacy depends on how a product is built and configured.

The reported 11.1 ms time to first token and high decoding rate point to a design intended for responsive transcription. These figures are useful indications, not a guarantee of performance on every device: Cactus measured them on an Apple M4 Pro and compared models through their respective official runtimes and default settings.

Whistle is also designed to work alongside Needle in a shared C++ engine. Cactus says this lets one binary handle audio transcription and then pass text toward tool calls. That could simplify an integrated on-device system, but the report does not provide independent tests of a complete speech-to-tool workflow or establish how it performs across different hardware.

Amazon

on-device speech recognition software

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

How Whistle Uses Needle

Cactus presents Whistle as a speech model built on components shared with Needle, its existing model. Its report says the eight-block encoder and decoder use Needle’s Simple Attention and Laddered Simple Attention blocks, with speech-specific gated cross-attention connecting the decoder to audio representations. The company says the shared blocks run Needle’s code rather than a separate copy.

The processing path begins by converting audio into log-mel features, then reduces the sequence to one frame per 80 milliseconds. The decoder uses five-beam search, and optional keyword biasing can raise the probability of supplied phrases. Cactus says decoder depth can be selected at load time, from two layers upward, while all eight encoder blocks still run.

The benchmark table is not a uniform head-to-head across every dataset. Cactus says Whistle performs better on LibriSpeech test-clean and test-other, SPGISpeech, Earnings-22 and the FLEURS average. Whisper base is ahead on TED-LIUM, AMI and the MLS average. The report also says some entries are unavailable because other model authors did not publish results; Whisper’s AMI score uses AMI-IHM, a different subset from the AMI scores reported for the other systems.

““It is one 16.9 MB file, runs on the CPU with no dependencies, and loads into the same C++ engine as Needle.””

— Cactus Compute, in its October 2 release

Amazon

compact speech-to-text app

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

Independent Tests Still Needed

The release is a vendor report, and the supplied material does not include independent replication of its accuracy, speed or memory claims. The stated speed results cover one Apple M4 Pro CPU and a 10-second audio clip; performance on phones, microcontrollers and other listed target devices is not detailed.

Cactus’s benchmark results also depend on dataset, runtime and evaluation conditions. Its report notes that Whisper’s AMI result comes from a different subset, while some models have no published score for certain tests. The supplied material does not give enough detail to independently assess all scoring and configuration choices.

The company describes Whistle as open, but the supplied report does not specify the license, model-weight access terms or full deployment requirements. It also does not quantify the model’s memory use while running, power consumption, transcription quality in noisy conditions or the limits of the silence threshold.

Amazon

voice transcription for wearables

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

Availability and Device Testing

Cactus says Whistle is released and provides a browser demonstration that downloads the 16.9 MB model for local use. The company’s report does not set out a later release date or identify a schedule for additional language support.

The next useful evidence will be documentation and testing beyond the company’s own comparison: licensing and installation details, results on representative target devices, and independent checks of accuracy across languages and audio conditions. Until those are available, the published figures describe Cactus’s reported results rather than a general guarantee of real-world performance.

Amazon

embedded speech recognition model

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

Key Questions

What is Whistle?

Whistle is an open speech-recognition model from Cactus Compute. The company says it transcribes audio locally using a CPU-based C++ engine shared with Needle.

Which languages does Whistle support?

Cactus lists English, German, French, Spanish, Italian, Dutch and Polish. It says the model detects the language unless a user names it.

Does Whistle send audio to the cloud?

Cactus says audio in its browser demonstration stays on the device after the model downloads. The report does not establish how every third-party product using Whistle will handle audio.

How does it compare with Whisper and Moonshine?

Cactus reports a smaller file size and faster time to first token for Whistle in its comparison. Accuracy results vary by dataset, and the tests used each model’s official runtime and default settings, so the figures should not be treated as independent, universal rankings.

What remains unknown about Whistle?

The supplied release does not establish performance across different devices, independent benchmark results, runtime memory use, power consumption or the model’s license terms. Those details matter for evaluating deployment in a specific product.

Source: hn

FALL

Fall Picks

As an affiliate, we earn on qualifying purchases.

You May Also Like

Why More Businesses Are Turning To Mistral Forge AI

An increasing number of organizations are adopting Mistral Forge AI for high-sovereignty, specialized AI needs, driven by strict data control and technical maturity requirements.

Chaos Engineering Techniques: Breaking Systems to Improve Resilience

Lifting the veil on chaos engineering techniques reveals how intentionally breaking systems can unlock greater resilience and reliability—discover the secrets inside.

Are You A Do-It-All Parent Or A Single Parent? Trends Shaping Parenthood Today

Exploring the evolving landscape of parenthood, this report examines whether parents are becoming more of do-it-all figures or if single parenting is on the rise, and why it matters.

MkLinux and the pimped-out Apple Workgroup Server 9150

Developers successfully port MkLinux to a heavily modified Apple Workgroup Server 9150, marking a milestone in open-source hardware adaptation.