Đa phương thứcASRNgười nói

SotaSpeech

Chuyển giọng nói thành văn bản có mốc thời gian và gán người nói cho tiếng Việt và âm thanh đa ngôn ngữ — nhận dạng, gán người nói và xác định thời gian trong một lượt duy nhất.

Dùng thử trực tiếpTài liệu API
Tác vụ
SATS
Ngữ cảnh
128k tokens
Ngôn ngữ
VI · EN · multi

Mô hình thực hiện đồng thời việc chuyển văn bản, gán người nói và dự đoán mốc thời gian, tạo ra các đoạn văn bản căn theo thời gian kèm mốc tuyệt đối trên toàn bản ghi. Mô hình được xây dựng cho biên bản họp, phân tích cuộc gọi và xử lý media dạng dài.

Chuyển văn bản đa người nói với độ chồng lấn cao

DoubaoElevenLabsGPT-4oGemini 2.5 ProGemini 3 Pro
SotaSpeech Transcribe Diarize
0510152025303510.211.69.11415.27.8WER(Word Error Rate)28.418.215.323.124.613.9cpWER(Concatenated Permutation)19.17.46.89.68.96DER(Diarization Error Rate)

So sánh benchmark — tập kiểm thử cuộc họp Việt Nam

Gemini-Flash-3.7Mai_transcribe_1.5 (Microsoft)Elevenlabs_scribe_v2SmartVoice VNPTViettelAI
SotaSpeech
010203040506028.8635.8535.1752.5648.367.8WER(Word Error Rate)20.7128.7629.8746.7138.485.94NWER(Normalized WER)19.3928.6627.9340.4533.444.83CER(Character Error Rate)

Tỷ lệ lỗi trên tập kiểm thử cuộc họp Việt Nam — càng thấp càng tốt.

Lỗi thấp hơn so với nhà cung cấp trong nước
ViettelAI84% WER86% CERSmartVoice VNPT85% WER88% CER

Kiến trúc

1 · ingestuploadffmpeg decode16 kHz mono f32VADFireRedVAD | RMSRequest gate4 concurrentsgl-omniT=0 · 8192 tokwhole clip, no chunking2 · MOSS-Transcribe-Diarize + Vietnamese LoRAWhisper-Medium audio encoder80 mel bins · 24 layers · d_model 1024 · 16 heads · FFN 4096log-mel in · GELU30 s window,1500 positions= 50 frames/s4x time merge(B, T, 1024) reshape to (B, T/4, 4096)50 to 12.5tokens/sVQAdaptor — projectionLinear 4096 to 1024 · SiLU · Linear 1024 to 1024LayerNorm eps 1e-6maps encoderdim to LMhidden 1024masked_scatter into text embeddingsaudio placeholder token 151671Qwen3-0.6B decoder28 layers · d 1024 · GQA 16/8 · head_dim 128 · SiLURoPE theta 1e6 · RMSNorm · 131072-token context+ Vietnamese LoRA r=64 · ckpt-27000tied embeddingslm_head to 151936vocabjointly: words+ speaker + timeone autoregressive token stream3 · parse[0.11][S01]Goodmorning![1.03][1.11][S02]Morning,guys![1.34]timestampspeakerwordsspeaker tag: base only — VI adapter emits noneParse[t] ([Sxx])? text [t]Collapse repeatsSegment[]no CAPU — modelemits cased text

Mô hình nền OpenMOSS-Team/MOSS-Transcribe-Diarize (Apache-2.0), kết hợp bộ mã hoá Whisper với bộ giải mã Qwen3; thiết kế ngữ cảnh dài theo Yu et al., arXiv:2601.01554v7 (CC BY 4.0). Adapter tiếng Việt và toàn bộ dịch vụ xung quanh là của riêng dự án này.

Khả năng của mô hình

Chuyển văn bản kèm tách người nói

Đang tải…