Voice messages, video notes, audio, video, recordings forwarded from other chats, and links. More than 90 languages.
Audio: MP3, M4A, WAV, OGG, OPUS, FLAC, AAC, AMR, WMA, AIFF, and 3GA.
Video: MP4, MOV, WebM, MKV, AVI, WMV, MPEG, and 3GP.
The file must contain an audio track. A damaged file or a rare unsupported codec may fail even if its extension appears in the list.
FREE without a purchased package accepts files up to 300 MB. START, PRO, POWER, and FREE while a purchased package still has transcription minutes accept files up to 2 GB.
Recording-duration limits are explained under “Plans, packages and payment”.
Yes. Send a voice message or video note directly, or forward it to the chat with @taktextbot. After successful processing, the bot returns the text.
The same transcription settings apply as for audio and video files. Learn more on the voice-message transcription page.
Yes. Tap “Forward” on a voice message, video note, audio, or video, then send the recording to TAK! TEXT.
The bot does not access the original chat—it processes only the forwarded recording. Whether the original sender’s details are visible depends on Telegram’s forwarding settings.
Send the bot one public link to a specific audio or video recording. If a message contains several links, the bot processes only the first one.
Examples of supported content:
• YouTube videos and Shorts;
• videos on RuTube and VK Video;
• TikTok videos and Instagram Reels;
• public X/Twitter posts containing video;
• individual podcast episodes, including Spotify when a public audio file is available;
• public files on Google Drive, Dropbox, and Yandex Disk.
Some other public platforms may also work. Profiles, channels, playlists, live streams, and recordings that require sign-in are not supported.
Link transcription is available on START, PRO, POWER, and on FREE while a purchased package still has transcription minutes. Whether a particular recording can be processed depends on the platform’s rules and technical restrictions.
The bot supports more than 90 languages. The exact set depends on the mode: Quality supports more languages than Speed.
By default, the language is detected automatically. If the bot gets it wrong, or you usually send recordings in one language, you can select the recognition language manually. The bot may switch modes if the current mode does not support the selected language.
Recognition language and interface language are configured separately.
“Speakers” separates the text into turns by different people and labels them “Speaker 1”, “Speaker 2”, and so on.
“Timestamps” adds time markers about every 30 seconds, making it easier to find the relevant point in the recording.
The settings are independent: you can enable either one or both. They are available on every plan, including FREE.
Speaker separation works best with clean audio where people do not talk over one another.
Reply to a voice message, video note, audio, or video and mention @taktextbot in the same reply.
You do not need to add the bot to the chat. If TAK! TEXT is added as a permanent group member, it leaves automatically.
In the original chat, the bot shows a short preview covering no more than the first 2 minutes. If you have already started the bot and the recording fits your plan, file-size, and available-balance limits, the full result arrives in your private chat.
Before first starting the bot, three trial transcriptions are available, each covering no more than the first 2 minutes.
The bot receives only the message that mentions it and the recording being replied to. It cannot see the rest of the conversation or the participant list.
Speed and Quality modes. 95%+ accuracy on clean speech. Switch with one tap.
"Speed" mode processes audio fast — short voice messages are ready in seconds. "Quality" mode is more accurate, supports more languages, and handles noisy audio, accents, and overlapping speech better, though it can take a little longer. Both modes are available on every plan and switch with one tap in the bot's settings.
On clean speech — above 95% in "Quality" mode and 92–95% in "Speed" mode. Accuracy drops with heavy background noise, low-quality recordings, and overlapping speech from multiple people. Language also matters — common languages score higher; for rare ones you can manually set the recording's language in settings.
In the bot, open the settings menu (the "⚙️ Settings" button or the /settings command). Choose "Recognition mode" and toggle between "Speed" and "Quality". The setting persists for all future transcriptions.
Open the bot menu → "Settings" → "Transcription settings" → "Recognition language". Pick a language from the list or leave "Auto-detect". The chosen language behaves differently depending on the mode:
In "Quality" mode it's a hint to the bot: it boosts accuracy for the selected language, but audio in other languages is still recognized. Highest accuracy = "Quality" + an explicitly chosen language.
In "Speed" mode it's a strict constraint: results come back as fast as possible, but audio in other languages won't be recognized.
Summary, translation into 13 languages, Q&A. Summary is free on every plan.
After transcription, three tools are available. Summary — a brief 3–5 bullet-point recap (free on every plan). Translation — into 13 languages. Ask the text (Q&A) — ask a question about the transcript and get an answer grounded in the text.
Yes. The brief recap (summary) is free on every plan, including FREE. Translation and Q&A use AI requests: 10 free per month on FREE, 50 on START, unlimited on PRO and POWER.
13 languages: English, Spanish, French, German, Portuguese, Italian, Arabic, Turkish, Persian, Russian, Ukrainian, Uzbek, and Kazakh. Translation is triggered by the "Translate" button after a transcript is ready. We're expanding the list — if a specific language is missing, reach out to support.
Ask a question about the transcript and the bot will answer based on the text. For example: "What did we agree on?", "When is the next meeting?", "Which numbers were mentioned?". It's handy for long recordings when you need to find specific information quickly. Questions must be about the transcript content — the bot won't answer off-topic questions.
Plans and one-time packages, payment methods, subscription management, and how long minutes remain valid.
The bot has the free FREE plan and the paid monthly START, PRO, and POWER plans.
FREE includes 30 transcription minutes during the first 30 days, then 15 minutes every 30 days, plus 10 AI requests. On FREE, recordings can be up to 5 minutes and files up to 300 MB.
START includes 300 minutes and 50 AI requests per month; PRO includes 1,000 minutes and unlimited AI requests; POWER includes 5,000 minutes and unlimited AI requests.
START, PRO, and POWER include link transcription, files up to 2 GB, and processing with no separate plan-specific recording-duration limit.
A one-time package of minutes and AI requests can be added on any plan. See the detailed comparison and current prices under “Plans and packages”.
Monthly plans can be paid through Stripe, Telegram Stars, or Tribute. One-time packages of minutes and AI requests can be paid through Stripe or Telegram Stars.
The Stripe page may offer Visa and Mastercard cards, PayPal, Apple Pay, Google Pay, Alipay, WeChat Pay, and regional payment methods. The exact set depends on your country, currency, and device.
Telegram Stars lets you pay inside Telegram. For subscriptions, Tribute supports bank cards, including MIR, Visa, and Mastercard, as well as payment through Wallet in Telegram.
Open /balance and tap “💳 Manage subscription”.
• For Stripe, open the customer portal, select your active TAK! TEXT subscription, and turn off renewal.
• For Telegram Stars, choose “🔕 Cancel auto-renewal” and confirm.
• For Tribute, open @tribute, go to “Subscriptions”, select TAK! TEXT, and tap “Cancel subscription”.
Only the next automatic renewal is disabled. Your paid plan remains available until the date shown in /balance.
A paid plan renews automatically every month until you turn off auto-renewal. Included minutes and AI requests refresh each paid period and do not carry over to the next one.
A package is paid for once, remains valid for 365 days, does not renew automatically, and does not change your current plan. Plan allowances are used first, followed by package allowances.
On FREE, while a purchased package still has transcription minutes, it also unlocks link transcription, files up to 2 GB, and processing with no separate plan-specific recording-duration limit.
On FREE without a purchased package, you can process up to 5 minutes of one recording. If the recording is longer, you can choose to transcribe its first 5 minutes.
START, PRO, POWER, and FREE with a purchased minute package have no separate plan-specific recording-duration limit.
Recordings up to several hours may be processed, but the actual maximum duration depends on file size, format, codec, and technical limits of the recognition service. A very long recording may take longer to process.
GDPR, servers in the EU. Audio is deleted immediately, transcripts after 24 hours.
Audio and video files are deleted immediately after processing — the bot doesn't store them. Transcripts are kept for up to 24 hours so the AI tools can work, then they're deleted automatically and permanently.
As a European company, we're required to comply with EU court orders. But since we don't store audio at all, and transcripts live no longer than 24 hours, in practice there's effectively nothing to hand over — even with an official request.
So that AI tools can run on the transcript. Summary, Q&A, and translation all run on the stored transcript. If it's already gone from the server, these features are technically impossible. The same goes for the "Download transcript" button — for long recordings that don't fit in a single Telegram message, downloading won't be available after 24 hours.
The transcript message itself stays in your chat with the bot — only you can see it, and we no longer have that data on our servers.
The core TAK! TEXT infrastructure runs on Hetzner Online GmbH, with data centers in Germany (EU). Certain processing operations may be performed by connected providers, including outside the EU, under appropriate safeguards (SCCs, DPA). Full list — in our privacy policy. More on security — on the dedicated page.
Yes. TAK! TEXT processes data in accordance with GDPR. The data controller is sershiko (Netherlands, KVK 42031706). You can request deletion of your data or a copy of it — email privacy@taktext.com. Details — in our privacy policy.
Transcripts are deleted automatically after 24 hours. Audio files are not stored — they're deleted immediately after processing.
If you need data deleted sooner or right away — the fastest path is to message the support bot. Official GDPR deletion requests are handled through privacy@taktext.com.