A CPU-oriented speech-recognition model using heterogeneous quantization: INT8 for its VAE and ternary I2 weights for the language model. It compresses the checkpoint from 4.62GB to 1.58GB, runs in real time on three CPU threads, and reports 1.6–2.3x higher speed than Whisper.cpp.

Model Details

License MIT

Paper

audiospeechefficientquantizationopen-weight