Updated: 2026-07-24
An agent's voice in Corpis is two independent abilities: speech recognition on input (you dictate a message instead of typing) and a spoken reply on output (the agent voices its answer with a character you set). Both work in the on-platform chat. To be honest: in a connected Telegram bot the agent answers with text — the voice abilities live in the on-platform chat.
Create an agent. Open agent creation and set the name, personality and greeting — or describe what you need in words and the AI assistant fills the form.
Open the "Voice" tab. In the agent settings go to "Voice" — this is where its sound profile lives.
Pick a preset and timbre. Take a ready preset (neutral, stern, cute, child, robot and others) and preview the timbres with a button — pick the one you like.
Fine-tune the character. Move the axes — pitch, rate, liveliness, energy, breathiness — or describe the voice in words, and the tab assistant picks and saves the profile.
Turn on spoken replies in chat. In the agent chat enable the speaker toggle in the input field — each new reply from the agent will be read out loud.
Dictate a message. Press the dictation button and say your message aloud — the speech is recognized into text and the agent replies to it.
Speech recognition (input) turns a dictated message into text: you speak, the agent gets the transcript and replies as it would to any message. This runs on the Whisper neural model right on the server — no cloud, no keys.
The spoken reply (output) is the reverse: the reply text is synthesized into voice and you hear it. Synthesis is local too (the Vosk-TTS engine), and the voice carries the character you gave the agent. These are two separate mechanisms and are turned on independently.
Voice — both recognition and spoken replies — works in the on-platform chat: there is a dictation button and a speaker toggle that reads out each new reply from the agent.
In a Telegram bot there is no voice yet: the agent answers with text, like an ordinary bot. This is an honest current limit, not a setting you missed — voice output in Telegram is planned but not built. So do not expect to "call" the bot in Telegram; voice currently lives in the on-platform chat.
Voice is set per agent, on the "Voice" tab. There are ready presets — neutral, stern, cute, child, cheerful, business, friendly, calm, robot — plus fine control with axes: pitch (higher is younger, lower is larger), rate, liveliness, energy, breathiness.
Both recognition and synthesis are neural models running on the Corpis server, not external paid APIs. There is no separate voice plan: the ability is available on the free plan too, and usage is counted against the agent's regular brain credits, same as a normal reply.
A "child" or "giant" voice is an approximation via pitch, not separately recorded child voices (the engines have none). Copying a specific person's voice from a sample is not possible right now — that is a class of heavy GPU models, a future tier; today the character comes from presets and axes. If synthesis fails, the reply simply stays as text in the thread — a robotic fallback voice is deliberately not switched in.
No: in Telegram the agent answers with text. Both speech recognition and spoken replies work in the on-platform chat, not in the Telegram bot.
No. Speech recognition and synthesis are neural models running right on the Corpis server, without external paid APIs or keys. A key is only optional, if you want to connect Yandex SpeechKit for acted emotion.
Yes. Each agent has its own voice profile on the "Voice" tab: nine presets (neutral, stern, cute, child, robot and others), five timbres, pitch/rate/liveliness axes, and "robot" and "over the radio" effects. You can also build the profile with words via the AI assistant.
No, cloning a voice from a reference is not available right now — that is a class of heavy GPU models and a future tier. The voice character is set with presets and tuning axes, and a "child" voice is an approximation via pitch rather than a separate recording.