Agentforce Platform | Speech to Text
Converts an audio file into text and returns the transcript for use in agent conversations or downstream actions.
Required Editions
| Available in: Lightning Experience |
| Available in: Enterprise, Performance, Unlimited, and Developer Editions with Foundations, or Agentforce 1 or Einstein 1 Editions |
| User Permissions Needed | |
|---|---|
| See Common User Access for Standard Agent Actions. | |
Action Details
| API Name | SpeechToText |
| Reference Action Type | Standard Action |
| Reference Action | Speech to Text |
| Does this action execute one or more prompt templates? | No |
Guidelines and Considerations
- The Speech to Text agent action isn’t supported in the Government Cloud.
- This action consumes Flex Credits under the Speech to Text usage type.
- This action returns a transcript only and doesn’t store it automatically. To save it, use Flow elements such as Query Records or Create Records, or Apex logic to write the output to a Salesforce record.
- This action supports multiple languages and common audio formats such as MP3 (MPEG Audio Layer III), WAV (Waveform Audio File Format), FLAC (Free Lossless Audio Codec), OGG or OGA (Ogg Vorbis), AMR (Adaptive Multi-Rate), MPEG (Moving Picture Experts Group), MPGA (MPEG Audio), with a maximum file size of 5 MB.
- This action transcribes audio in the primary detected language, and mixed-language audio may work. If you plan to use the transcript in Agentforce downstream features, verify that the detected language is supported. See Generative AI Supported Languages.
- This action does not perform translation.
For detailed usage guidelines and limitations, see Considerations and Limitations for Speech to Text in Agentforce.
Did this article solve your issue?
Let us know so we can improve!

