Generate Speech

cURL

curl --request POST \
  --url https://api.heygen.com/v3/voices/speech \
  --header 'Content-Type: application/json' \
  --header 'x-api-key: <api-key>' \
  --data '\n{\n  "text": "<string>",\n  "voice_id": "<string>",\n  "input_type": "text",\n  "speed": 1,\n  "language": "<string>",\n  "locale": "<string>"\n}\n'

Possible Status Codes

  • 200
  • 400
  • 401
  • 429
{
  "data": {
    "audio_url": "<string>",
    "duration": 123,
    "request_id": "req_abc123",
    "word_timestamps": [
      {
        "word": "<string>",
        "start": 123,
        "end": 123
      }
    ]
  }
}

Authorizations

ApiKeyAuth

  • x-api-key
    Type: string
    Location: header
    Required: Yes
    Description: HeyGen API key. Obtain from your HeyGen dashboard.

Body

application/json

  • text
    Type: string
    Required: Yes
    Description: Text to synthesize (1-5000 characters).
    Required string length: 1 - 5000
  • voice_id
    Type: string
    Required: Yes
    Description: Voice ID to use. The voice must support the starfish engine. Filter compatible voices by passing engine=starfish to the voice listing endpoint.
  • input_type
    Type: string
    Default: text
    Description: Type of the input: 'text' for plain text, 'ssml' for SSML markup. Defaults to 'text'.
  • speed
    Type: number
    Default: 1
    Description: Speed multiplier (0.5-2.0).
    Required range: 0.5 <= x <= 2
  • language
    Type: string | null
    Description: Base language code (e.g. 'en', 'pt', 'zh'). Optional — auto-detected from text when omitted.
  • locale
    Type: string | null
    Description: BCP-47 locale tag (e.g. 'en-US', 'pt-BR'). When set, language is inferred from locale.

Response

200

application/json

  • data
    Description: Response payload for text-to-speech generation.