Kotoba Technologies is a frontier voice AI company building a real-time speech model and simultaneous translation technology, with a focus on exceptional performance in East Asian languages. The company's core product is Koto, an end-to-end voice AI model built for ultra-low-latency, real-time speech that delivers speech-to-text, text-to-speech, and speech-to-speech translation capabilities. Koto achieves sub-50ms TTS latency, ultra-fast STT, and speech-to-speech translation at the speed of a professional simultaneous interpreter. The model runs flexibly across on-premise, on-cloud, and on-device deployments, making it suitable for AI agents, real-time interpretation devices, smart glasses, and mobile applications. Kotoba also offers a consumer-facing app providing real-time simultaneous translation so conversations can flow naturally across languages, delivering fast and accurate translation in both voice and text. The company provides an API and SDK (in alpha) offering speech-to-speech, streaming STT, and TTS for builders. Kotoba has partnered with jig.jp to bring real-time AI interpretation to SABERA smart glasses. The company differentiates itself through its specialized focus on East Asian languages where most general-purpose voice AI models underperform, its ultra-low-latency architecture enabling true conversational-speed translation, and its flexible deployment model supporting on-device inference for privacy-sensitive applications. Founded in 2023 and headquartered in the United States, Kotoba Technologies has raised approximately $23 million in a Seed round (including a $10 million seed funding round announced in June 2026) backed by Globis Capital Partners, Boostcapital, SIP Capital, Kindred Ventures, Salesforce Ventures, and Sony Innovation Fund.
Kotoba Technologies
SeedKotoba Technologies is a frontier voice AI company building a real-time speech model and simultaneous translation technology, with a focus on exceptional performance in East Asian languages. The company's core product is Koto, an end-to-end voice AI model built for ultra-low-latency, real-time speech that delivers speech-to-text, text-to-speech, and speech-to-speech translation capabilities. Koto achieves sub-50ms…
About
Business model
Speech and Voice Recognition > Technology Providers > Software > Automatic Speech Recognition,
AI Infrastructure > Natural Language Processing > Speech Solutions > Speech Recognition
Team background
College Wise > University of Michigan-Ann Arbor, Cornell University, Yale University
Institutional investors
Latest coverage
📰 See all funding news — keep up with the raises that set comps.
Backed by funds in our directory
🗂️ Full list of active funds — filter by stage, check size, and geography.