Ganas Speech to Text — developer integration guide

Base URL: https://api.stt.ganas.ai
Alternative hostname: https://www.api.stt.ganas.ai
Live WebSocket: wss://api.stt.ganas.ai/v1/audio/transcriptions/ws

This deployment transcribes English, Hindi, and Marathi, with automatic language selection. It produces finalized utterance transcripts; it does not currently emit partial word-by-word hypotheses, speaker diarization, word timestamps, or translations.

Authentication and access

Use the client API key issued to your application. For HTTP, send Authorization: Bearer CLIENT_KEY. For WebSocket, send {"api_key":"CLIENT_KEY"} as the first text frame within five seconds, or provide the same Authorization header during connection setup. Do not put keys in URLs. Wait for the ready event before sending audio.

Keys have expiration, revocation, source-IP rules, request-rate limits, and simultaneous-request limits. HTTP uploads and open authenticated WebSocket sessions count toward the simultaneous-request limit. Rate limits count HTTP requests and new WebSocket sessions; they do not count individual audio frames. A global IP rule and the key's IP rule must both permit the caller. Behind NAT, register the application's public egress IP.

Keep long-lived keys on your backend. For a browser integration on a different website, arrange an approved origin and proxy requests through your backend; do not embed a shared client secret in publicly served JavaScript.

Endpoints

Method Path Purpose Authentication
GET /health Process liveness None
GET /ready Readiness and configured languages None
POST /v1/audio/transcriptions Transcribe an uploaded audio file Bearer key
WebSocket /v1/audio/transcriptions/ws Continuous PCM audio; finalized utterances First JSON frame or Bearer header
WebSocket /ws/telephony Telephony mu-law audio ingestion First JSON frame or Bearer header
GET /calls/{call_id}/transcripts Retrieve your key's telephony transcripts Bearer key

POST /transcribe and WebSocket /ws/transcribe remain supported aliases. The file API uses multipart upload; it is not a drop-in implementation of every field in other vendors' transcription APIs.

File transcription

Upload a multipart/form-data field named file. Supported inputs include WAV, MP3, FLAC and OGG. Audio is decoded to mono 16 kHz. The file limit is 25 MiB, and the total multipart body limit is 26 MiB. Clips longer than 30 seconds are rejected: split them or use the continuous WebSocket API.

export STT_API_KEY='YOUR_CLIENT_KEY'
curl --fail-with-body 'https://api.stt.ganas.ai/v1/audio/transcriptions' \
  -H "Authorization: Bearer $STT_API_KEY" \
  -F 'file=@sample.wav'

Optional query parameters:

Parameter Default Range Meaning
max_new_tokens 256 1–512 Output generation cap; low values can cut off the transcript
repetition_penalty 1.2 1.0–2.0 Repeated-output control

Example JSON response (values are illustrative):

{
  "text": "Hello, how can I help you?",
  "duration_seconds": 1.25,
  "audio_seconds": 3.8,
  "truncated": false,
  "avg_logprob": -0.18,
  "is_low_confidence": false,
  "is_likely_hallucination": false,
  "is_disallowed_script": false
}

duration_seconds measures inference plus queue waiting, not the entire upload time. audio_seconds is the decoded clip duration. avg_logprob is an internal confidence signal, not an accuracy percentage; closer to zero generally indicates more confident output. Treat the confidence/repetition flags as reasons to review the result. A disallowed-script result has empty text. Silence or noisy input can still produce incorrect words; evaluate with your real recordings.

Python (pip install requests):

import os
import requests

with open("sample.wav", "rb") as audio:
    response = requests.post(
        "https://api.stt.ganas.ai/v1/audio/transcriptions",
        headers={"Authorization": f"Bearer {os.environ['STT_API_KEY']}"},
        files={"file": ("sample.wav", audio, "audio/wav")},
        timeout=(10, 120),
    )
response.raise_for_status()
print(response.json()["text"])

Node.js 20+:

import { readFile } from 'node:fs/promises';
const form = new FormData();
form.append('file', new Blob([await readFile('sample.wav')], {type:'audio/wav'}), 'sample.wav');
const response = await fetch('https://api.stt.ganas.ai/v1/audio/transcriptions', {
  method:'POST', headers:{Authorization:`Bearer ${process.env.STT_API_KEY}`}, body:form
});
if (!response.ok) throw new Error(`${response.status}: ${await response.text()}`);
console.log(await response.json());

Let your HTTP library set the multipart Content-Type and boundary automatically.

Live WebSocket transcription

Connect, authenticate, wait for ready, then send binary audio frames containing signed 16-bit little-endian PCM, mono, 16,000 Hz, with no WAV/MP3 header. Recommended chunks: 20–100 ms (640–3,200 bytes). Each binary frame must have an even byte count and be at most 65,536 bytes. Send audio at approximately its real-time rate.

Handshake:

{"api_key":"YOUR_CLIENT_KEY"}

Server:

{"event":"ready","format":"pcm_s16le","sample_rate":16000,"channels":1}

The server detects utterances and sends a finalized transcript event after a pause (configured silence threshold: approximately 700 ms). Continuous speech is divided at approximately 28 seconds. Audio buffering, queueing and inference add latency; there is no fixed latency guarantee.

{
  "event":"transcript",
  "text":"नमस्ते, आप कैसे हैं?",
  "audio_seconds":2.4,
  "inference_seconds":0.8,
  "queue_wait_seconds":0.0,
  "avg_logprob":-0.2,
  "is_low_confidence":false,
  "is_likely_hallucination":false,
  "is_disallowed_script":false,
  "truncated_by_cap":false
}

Client text controls:

Message Behaviour
{"event":"end"} Flush the current utterance; receive transcript event(s), possibly empty, followed by {"event":"flushed"}. Connection remains open.
{"event":"reset"} Discard the pending utterance; receive {"event":"reset_ack"}.
{"event":"ping"} Receive {"event":"pong"}.

No incoming audio/control message for 30 seconds closes an idle connection. Send ping during intentional pauses. Before ending a recording, send end, wait for flushed, then close the socket. Reconnect with backoff if the connection drops; there is no session resumption or replay deduplication.

Python (pip install websockets), using a PCM16 mono 16 kHz WAV source:

import asyncio, json, os, wave
from websockets.asyncio.client import connect

async def main():
    async with connect("wss://api.stt.ganas.ai/v1/audio/transcriptions/ws") as ws:
        await ws.send(json.dumps({"api_key": os.environ["STT_API_KEY"]}))
        ready = json.loads(await ws.recv())
        if ready.get("event") != "ready":
            raise RuntimeError(ready)

        flushed = asyncio.Event()

        async def receive():
            async for message in ws:
                event = json.loads(message)
                if event.get("event") == "error":
                    raise RuntimeError(event)
                if event.get("event") == "flushed":
                    flushed.set()
                if event.get("event") == "transcript":
                    print(event.get("text", ""))

        receiver = asyncio.create_task(receive())
        try:
            with wave.open("sample.wav", "rb") as audio:
                assert (audio.getnchannels(), audio.getsampwidth(), audio.getframerate()) == (1, 2, 16000)
                while chunk := audio.readframes(800):
                    await ws.send(chunk)
                    await asyncio.sleep(0.05)
            await ws.send(json.dumps({"event": "end"}))
            await asyncio.wait_for(flushed.wait(), timeout=120)
        finally:
            await ws.close()
            await receiver

asyncio.run(main())

Telephony WebSocket

Use wss://api.stt.ganas.ai/ws/telephony. Authenticate and wait for ready just as above. Then send JSON start, media, and stop events. Audio must be G.711 mu-law, mono, 8,000 Hz. This is a generic adapter protocol; confirm your telephony provider supports this authentication and frame shape, or use your own bridging service.

{"event":"start","start":{"streamSid":"stream_123","callSid":"call_123","mediaFormat":{"encoding":"audio/x-mulaw","sampleRate":8000,"channels":1}}}
{"event":"media","media":{"payload":"BASE64_MULAW_BYTES"}}
{"event":"stop"}

Call and stream IDs must use 1–128 ASCII letters, digits, underscores or hyphens. Use unique call IDs. Maximum JSON frame size is 65,536 bytes including base64. This endpoint stores finalized results for polling; it does not emit transcript events on the media socket. After stop, pending speech is finalized before the connection closes.

curl --fail-with-body 'https://api.stt.ganas.ai/calls/call_123/transcripts' \
  -H "Authorization: Bearer $STT_API_KEY"
{"call_id":"call_123","transcripts":[{"call_id":"call_123","stream_sid":"stream_123","call_sid":"call_123","text":"Hello","audio_seconds":1.2,"inference_seconds":0.8,"avg_logprob":-0.2,"is_low_confidence":false,"is_likely_hallucination":false,"is_disallowed_script":false,"truncated_by_cap":false,"ts":1790755200.0}]}

Polling returns a snapshot; it does not consume the records. Results are isolated by client key. Keep using the same key for ingestion and polling. Up to 200 utterances per call are retained temporarily in memory, expiring about one hour after the last result. Capacity eviction or a service restart can remove results earlier. Persist transcripts in your own application. Client-supplied webhook URLs are not supported.

Errors and capacity

HTTP errors contain detail (a string or a validation-error array). Proxy-generated errors may be HTML, so check the response Content-Type before parsing JSON. WebSocket errors use {"event":"error","detail":"...","status":401}; status is optional. Authentication/policy failures close with code 1008; unexpected service failures may use 1011. A rejected handshake may instead appear as HTTP 403.

Status Meaning Action
400 Empty, corrupt or undecodable audio Correct the upload
401 Missing, expired, revoked or invalid key Check your key
403 Source IP or browser origin denied Request approved access
413 Body or frame too large Reduce size
422 Invalid parameter or file longer than 30 seconds Fix fields or use streaming
429 Per-key rate/concurrency or global connection limit Retry with exponential backoff and jitter
503 Loading or full inference queue Retry with backoff
500/502/504 Service/proxy failure Retry boundedly and report persistent errors

This instance admits at most 8 active authenticated requests/sessions globally, further limited by your key. One GPU inference runs at a time, with a bounded queue. Eight connected clients does not imply eight simultaneous GPU decodes or a guaranteed throughput. Measure latency with your expected audio and traffic mix; STT shares server resources with other workloads.

Use separate connection/read timeouts, handle disconnects, and avoid unlimited retry loops. Reports record usage metadata; audio uploads are temporary, and live transcripts are returned to the caller. Telephony results use the temporary retention described above.