Generate content (Gemini)
Audio
Native Gemini Format
Gemini-native generateContent interface for text chat, multimodal media recognition (images, audio, video), speech synthesis, and image generation with structured parts. Use generationConfig to request specific response modalities such as speech (speechConfig) or images (imageConfig).
POST
Generate content (Gemini)
This page uses the same
generateContent operation as Generate content (Gemini), with the playground above pre-filled for plain text chat. The notes below describe the Gemini-native fields you can add to generationConfig to request audio understanding or generation with structured parts.
Set
generationConfig.responseModalities to ["AUDIO"] to request audio output, and configure generationConfig.speechConfig.voiceConfig.prebuiltVoiceConfig.voiceName to choose a prebuilt voice for generated speech.Gemini-native request fields
Example: requesting speech audio
Response fields
The response follows the standardgenerateContent shape. When audio output is requested, the returned parts contain inline audio data instead of text:
array
Candidate responses returned by the model.
object
Token accounting, including
promptTokenCount, candidatesTokenCount, and totalTokenCount.object
Prompt blocking feedback when applicable.
Example response
200
Authorizations
Your DGrid API key. All endpoints use Authorization: Bearer <DGRID_API_KEY>.
Path Parameters
Target model ID, such as gemini-1.5-pro.
Body
application/json

