Meralion and ST Engineering logo
Model Capabilities
Merlion and ST Engineering

Consortium Tuned Model (CTM) -
South East Asia's Leading Automatic Speech Recognition Model

Natively Built for Southeast Asia

Trained on Singapore and Southeast Asian speech data — not adapted from a Western model.

Consortium-Grade Reliability

Developed through A*STAR's MERaLiON consortium, with real-world validation from Grab, DBS, HTX, and more.

Built for Integration

API-first. Whether you're licensing for enterprise or building your own product, we have a pathway for you.

Government-Grade Security

Developed by ST Engineering with public sector deployability in mind.

ASR-LLM with 3 Advanced Capabilities.

MERaLiON × STE layers advanced speech intelligence on top of SEA's most culturally-aware foundation model.

01 — Live Transcription

Real-time ASR

Sub-300ms latency. Streams text as people speak — not after they finish.

<300ms

02 — Speaker Identification

N-way Diarisation

Labels every utterance by speaker. Works across 2 to N participants simultaneously.

N-speaker

03 — Clean Output

Contextualised Normalisation

Handles numbers, dates, proper nouns, and local idioms — no messy post-processing needed.

Ready-to-use

Speech AI that knows who said what.

Real-time transcription and n-way speaker diarization — built for Southeast Asia's languages, dialects, and multi-speaker environments.

Live diarization — 3 speakers detected

Speaker A

diarized audio waveform

Speaker B

diarized audio waveform

Speaker C

diarized audio waveform
Latency<300ms
LanguagesEN · ZH · MS · TA · Singlish

Ready to hear the difference?

See real-time diarization in action with your own audio. Book a 30-minute demo.

© 2026 ST Engineering. All rights reserved.