JEV-Speech: Same 24-Layer Encoder, 2.18x Faster Inference

JEV-Speech is an Orukeet offshoot runtime from Oruk Labs that returns a transcript plus auxiliary non-transcript outputs in less than half the time. Warm p95 request latency drops from 91.93 to 42.23 ms on an A100 at batch 1, a 2.18x speedup measured from a decoded waveform in host memory to completed outputs back on the host, with all 24 encoder layers intact. The honest part is the trade-off table: a variant that changes neither weights nor precision reaches 1.66x with every transcript unchanged across 250 recordings, while the faster candidate adds BF16 arithmetic and encoder adaptation and made fifteen more word errors than the original on a separate seven-language test. Measurements are warm batch 1 over 250 historical clips with six balanced passes.





