# Ganas Speech to Text — developer integration guide

Base URL: **https://api.stt.ganas.ai**  
Alternative hostname: **https://www.api.stt.ganas.ai**  
Live WebSocket: **wss://api.stt.ganas.ai/v1/audio/transcriptions/ws**

This deployment transcribes **English, Hindi, and Marathi**, with automatic language selection. It produces finalized utterance transcripts; it does not currently emit partial word-by-word hypotheses, speaker diarization, word timestamps, or translations.

## Authentication and access

Use the client API key issued to your application. For HTTP, send `Authorization: Bearer CLIENT_KEY`. For WebSocket, send `{"api_key":"CLIENT_KEY"}` as the first text frame within five seconds, or provide the same Authorization header during connection setup. Do not put keys in URLs. Wait for the `ready` event before sending audio.

Keys have expiration, revocation, source-IP rules, request-rate limits, and simultaneous-request limits. HTTP uploads and open authenticated WebSocket sessions count toward the simultaneous-request limit. Rate limits count HTTP requests and new WebSocket sessions; they do not count individual audio frames. A global IP rule and the key's IP rule must both permit the caller. Behind NAT, register the application's public egress IP.

Keep long-lived keys on your backend. For a browser integration on a different website, arrange an approved origin and proxy requests through your backend; do not embed a shared client secret in publicly served JavaScript.

## Endpoints

| Method | Path | Purpose | Authentication |
|---|---|---|---|
| GET | `/health` | Process liveness | None |
| GET | `/ready` | Readiness and configured languages | None |
| POST | `/v1/audio/transcriptions` | Transcribe an uploaded audio file | Bearer key |
| WebSocket | `/v1/audio/transcriptions/ws` | Continuous PCM audio; finalized utterances | First JSON frame or Bearer header |
| WebSocket | `/ws/telephony` | Telephony mu-law audio ingestion | First JSON frame or Bearer header |
| GET | `/calls/{call_id}/transcripts` | Retrieve your key's telephony transcripts | Bearer key |

`POST /transcribe` and WebSocket `/ws/transcribe` remain supported aliases. The file API uses multipart upload; it is not a drop-in implementation of every field in other vendors' transcription APIs.

## File transcription

Upload a `multipart/form-data` field named `file`. Supported inputs include WAV, MP3, FLAC and OGG. Audio is decoded to mono 16 kHz. The file limit is **25 MiB**, and the total multipart body limit is **26 MiB**. Clips longer than **30 seconds** are rejected: split them or use the continuous WebSocket API.

```bash
export STT_API_KEY='YOUR_CLIENT_KEY'
curl --fail-with-body 'https://api.stt.ganas.ai/v1/audio/transcriptions' \
  -H "Authorization: Bearer $STT_API_KEY" \
  -F 'file=@sample.wav'
```

Optional query parameters:

| Parameter | Default | Range | Meaning |
|---|---|---|---|
| `max_new_tokens` | 256 | 1–512 | Output generation cap; low values can cut off the transcript |
| `repetition_penalty` | 1.2 | 1.0–2.0 | Repeated-output control |

Example JSON response (values are illustrative):

```json
{
  "text": "Hello, how can I help you?",
  "duration_seconds": 1.25,
  "audio_seconds": 3.8,
  "truncated": false,
  "avg_logprob": -0.18,
  "is_low_confidence": false,
  "is_likely_hallucination": false,
  "is_disallowed_script": false
}
```

`duration_seconds` measures inference plus queue waiting, not the entire upload time. `audio_seconds` is the decoded clip duration. `avg_logprob` is an internal confidence signal, not an accuracy percentage; closer to zero generally indicates more confident output. Treat the confidence/repetition flags as reasons to review the result. A disallowed-script result has empty `text`. Silence or noisy input can still produce incorrect words; evaluate with your real recordings.

Python (`pip install requests`):

```python
import os
import requests

with open("sample.wav", "rb") as audio:
    response = requests.post(
        "https://api.stt.ganas.ai/v1/audio/transcriptions",
        headers={"Authorization": f"Bearer {os.environ['STT_API_KEY']}"},
        files={"file": ("sample.wav", audio, "audio/wav")},
        timeout=(10, 120),
    )
response.raise_for_status()
print(response.json()["text"])
```

Node.js 20+:

```javascript
import { readFile } from 'node:fs/promises';
const form = new FormData();
form.append('file', new Blob([await readFile('sample.wav')], {type:'audio/wav'}), 'sample.wav');
const response = await fetch('https://api.stt.ganas.ai/v1/audio/transcriptions', {
  method:'POST', headers:{Authorization:`Bearer ${process.env.STT_API_KEY}`}, body:form
});
if (!response.ok) throw new Error(`${response.status}: ${await response.text()}`);
console.log(await response.json());
```

Let your HTTP library set the multipart Content-Type and boundary automatically.

## Live WebSocket transcription

Connect, authenticate, wait for `ready`, then send binary audio frames containing **signed 16-bit little-endian PCM, mono, 16,000 Hz**, with no WAV/MP3 header. Recommended chunks: 20–100 ms (**640–3,200 bytes**). Each binary frame must have an even byte count and be at most **65,536 bytes**. Send audio at approximately its real-time rate.

Handshake:

```json
{"api_key":"YOUR_CLIENT_KEY"}
```

Server:

```json
{"event":"ready","format":"pcm_s16le","sample_rate":16000,"channels":1}
```

The server detects utterances and sends a finalized `transcript` event after a pause (configured silence threshold: approximately 700 ms). Continuous speech is divided at approximately 28 seconds. Audio buffering, queueing and inference add latency; there is no fixed latency guarantee.

```json
{
  "event":"transcript",
  "text":"नमस्ते, आप कैसे हैं?",
  "audio_seconds":2.4,
  "inference_seconds":0.8,
  "queue_wait_seconds":0.0,
  "avg_logprob":-0.2,
  "is_low_confidence":false,
  "is_likely_hallucination":false,
  "is_disallowed_script":false,
  "truncated_by_cap":false
}
```

Client text controls:

| Message | Behaviour |
|---|---|
| `{"event":"end"}` | Flush the current utterance; receive transcript event(s), possibly empty, followed by `{"event":"flushed"}`. Connection remains open. |
| `{"event":"reset"}` | Discard the pending utterance; receive `{"event":"reset_ack"}`. |
| `{"event":"ping"}` | Receive `{"event":"pong"}`. |

No incoming audio/control message for **30 seconds** closes an idle connection. Send `ping` during intentional pauses. Before ending a recording, send `end`, wait for `flushed`, then close the socket. Reconnect with backoff if the connection drops; there is no session resumption or replay deduplication.

Python (`pip install websockets`), using a PCM16 mono 16 kHz WAV source:

```python
import asyncio, json, os, wave
from websockets.asyncio.client import connect

async def main():
    async with connect("wss://api.stt.ganas.ai/v1/audio/transcriptions/ws") as ws:
        await ws.send(json.dumps({"api_key": os.environ["STT_API_KEY"]}))
        ready = json.loads(await ws.recv())
        if ready.get("event") != "ready":
            raise RuntimeError(ready)

        flushed = asyncio.Event()

        async def receive():
            async for message in ws:
                event = json.loads(message)
                if event.get("event") == "error":
                    raise RuntimeError(event)
                if event.get("event") == "flushed":
                    flushed.set()
                if event.get("event") == "transcript":
                    print(event.get("text", ""))

        receiver = asyncio.create_task(receive())
        try:
            with wave.open("sample.wav", "rb") as audio:
                assert (audio.getnchannels(), audio.getsampwidth(), audio.getframerate()) == (1, 2, 16000)
                while chunk := audio.readframes(800):
                    await ws.send(chunk)
                    await asyncio.sleep(0.05)
            await ws.send(json.dumps({"event": "end"}))
            await asyncio.wait_for(flushed.wait(), timeout=120)
        finally:
            await ws.close()
            await receiver

asyncio.run(main())
```

## Telephony WebSocket

Use `wss://api.stt.ganas.ai/ws/telephony`. Authenticate and wait for `ready` just as above. Then send JSON `start`, `media`, and `stop` events. Audio must be **G.711 mu-law, mono, 8,000 Hz**. This is a generic adapter protocol; confirm your telephony provider supports this authentication and frame shape, or use your own bridging service.

```json
{"event":"start","start":{"streamSid":"stream_123","callSid":"call_123","mediaFormat":{"encoding":"audio/x-mulaw","sampleRate":8000,"channels":1}}}
```

```json
{"event":"media","media":{"payload":"BASE64_MULAW_BYTES"}}
```

```json
{"event":"stop"}
```

Call and stream IDs must use 1–128 ASCII letters, digits, underscores or hyphens. Use unique call IDs. Maximum JSON frame size is 65,536 bytes including base64. This endpoint stores finalized results for polling; it does not emit transcript events on the media socket. After `stop`, pending speech is finalized before the connection closes.

```bash
curl --fail-with-body 'https://api.stt.ganas.ai/calls/call_123/transcripts' \
  -H "Authorization: Bearer $STT_API_KEY"
```

```json
{"call_id":"call_123","transcripts":[{"call_id":"call_123","stream_sid":"stream_123","call_sid":"call_123","text":"Hello","audio_seconds":1.2,"inference_seconds":0.8,"avg_logprob":-0.2,"is_low_confidence":false,"is_likely_hallucination":false,"is_disallowed_script":false,"truncated_by_cap":false,"ts":1790755200.0}]}
```

Polling returns a snapshot; it does not consume the records. Results are isolated by client key. Keep using the same key for ingestion and polling. Up to 200 utterances per call are retained temporarily in memory, expiring about one hour after the last result. Capacity eviction or a service restart can remove results earlier. Persist transcripts in your own application. Client-supplied webhook URLs are not supported.

## Errors and capacity

HTTP errors contain `detail` (a string or a validation-error array). Proxy-generated errors may be HTML, so check the response Content-Type before parsing JSON. WebSocket errors use `{"event":"error","detail":"...","status":401}`; `status` is optional. Authentication/policy failures close with code 1008; unexpected service failures may use 1011. A rejected handshake may instead appear as HTTP 403.

| Status | Meaning | Action |
|---|---|---|
| 400 | Empty, corrupt or undecodable audio | Correct the upload |
| 401 | Missing, expired, revoked or invalid key | Check your key |
| 403 | Source IP or browser origin denied | Request approved access |
| 413 | Body or frame too large | Reduce size |
| 422 | Invalid parameter or file longer than 30 seconds | Fix fields or use streaming |
| 429 | Per-key rate/concurrency or global connection limit | Retry with exponential backoff and jitter |
| 503 | Loading or full inference queue | Retry with backoff |
| 500/502/504 | Service/proxy failure | Retry boundedly and report persistent errors |

This instance admits at most **8 active authenticated requests/sessions** globally, further limited by your key. One GPU inference runs at a time, with a bounded queue. Eight connected clients does not imply eight simultaneous GPU decodes or a guaranteed throughput. Measure latency with your expected audio and traffic mix; STT shares server resources with other workloads.

Use separate connection/read timeouts, handle disconnects, and avoid unlimited retry loops. Reports record usage metadata; audio uploads are temporary, and live transcripts are returned to the caller. Telephony results use the temporary retention described above.
