Reference17:Apps/PbxManager/App Tts
With the text-to-speech (TTS) settings plugin, the required app object can be created and configured. In addition, the app object can be assigned to specific config templates, if any exist.
The TTS app has no user interface of its own. It provides the app platform API com.innovaphone.speech to other apps, forwards their synthesis requests to a remote text-to-speech provider and streams the resulting audio back to them. Everything an administrator has to set up is therefore the connection to that provider.
Configuration
Accounts
Up to three providers can be configured side by side, for example a hosted service and a self-hosted one, or two services with different voices. An account consists of the four values Account Name, API, Remote service URL and API key. All other settings on this page are shared by all accounts.
- Accounts
- Selects the account that is edited below. Switching the selection loads that account's API, Remote service URL and API key into the fields underneath, and - as soon as that account has a usable URL - reloads the model and voice lists from its provider.
- The first account in the list is the default account. It is used whenever an app requests speech without naming an account, so this should be the provider you want everybody to use.
- Account Name
- The name of the account currently selected, as apps see it. Changing it renames the entry in the Accounts list.
- Be Aware: Apps store the account they use by name. If you rename an account that is already in use, the apps fall back to the default account until the new name is selected there again. Names must be unique; an already used name is rejected and the previous name is restored.
Accounts two and three are only offered to the apps once their Remote service URL is filled in. An account with an empty URL stays invisible, so unused accounts do not have to be hidden in any other way.
Provider connection
The following three settings belong to the account selected in Accounts above.
- API
- The dialect the provider speaks - see Supported APIs below. This describes the shape of the interface, not the vendor: a service that reproduces the ElevenLabs interface is reached with Elevenlabs compatible even if it is operated by somebody else.
- Remote service URL
- The address of the remote text-to-speech service. Both the base address (
https://api.openai.com) and a full endpoint address (https://api.openai.com/v1/audio/speech) are accepted - everything from/v1onwards is replaced by the endpoint the selected API needs. - A path in front of the endpoint is kept, so services that are mounted below a prefix can be reached as well:
https://api.provider.com/11labsresults in requests tohttps://api.provider.com/11labs/v1/text-to-speech/….
- API key
- The secret key the provider expects, sent in the header the selected API prescribes. It is stored encrypted and is never sent back to the browser, which is why the field shows
***instead of the stored key. Leave the placeholder untouched to keep the key that is already stored; type a new key to replace it. - Services that do not require authentication - a self-hosted model, or a Gradio app - are used with an empty key.
Supported APIs
| API | Typical provider | Endpoint used for synthesis | Authentication | Where the Voice and LLM (model) lists come from |
|---|---|---|---|---|
| OpenAI compatible (default) | OpenAI, and self-hosted servers that reproduce its interface (Kokoro, LocalAI, …) | POST /v1/audio/speech
|
Authorization: Bearer {API key}
|
GET /v1/audio/voices and GET /v1/models. OpenAI itself publishes no voice list, therefore a built-in list of the OpenAI voices (alloy, ash, cedar, coral, echo, fable, marin, nova, onyx, shimmer) is offered whenever the provider returns none.
|
| Elevenlabs compatible | ElevenLabs, and proxies that reproduce its interface | POST /v1/text-to-speech/{voice}
|
xi-api-key: {API key}
|
GET /v1/voices and GET /v1/models. The voices are taken from the provider, including the voice samples used by the loudspeaker button next to Voice.
|
| Gradio compatible | A Gradio app that publishes a generate_speech endpoint
|
Three steps against /gradio_api/: the call, the event stream with the result, and the download of the generated file
|
none | GET /gradio_api/info, i.e. from the selection lists the Gradio app itself declares for its voice_name and model_choice parameters.
|
Speech settings
These settings are not part of an account; there is one set of them for the whole app.
- LLM (model)
- The model of the remote service that is used for the test playback below. The list is read from the provider of the selected account and is empty for providers that publish no models.
- Voice
- The voice used for the test playback below. The list is read from the provider of the selected account as described in Supported APIs.
- If the provider delivers a sample for the selected voice, a loudspeaker button appears next to the list which plays that sample. This is the provider's own sample and does not generate any speech.
- Extra instruction
- An additional instruction on how the text should be spoken, for example speak slowly and friendly. Only the OpenAI compatible API knows this parameter; the other two ignore it.
- Seed (-1 = random)
- Makes the synthesis repeatable, so that the same text is spoken the same way every time it is generated - which matters for announcements that are rendered again later. The value is passed on to the provider, which is only supported by the Elevenlabs compatible API; the other two ignore it.
-1, the default, sends no value and lets the provider choose freely. Permitted values are0to4294967295. Apps may send a seed of their own, which then takes precedence.
Be Aware: Voice and Extra instruction apply to the test playback on this page. Apps that use the speech service choose model and voice themselves Seed is the exception: it is used for every request that does not bring its own seed.
Testing the configuration
- Test phrase
- The text that is spoken when the loudspeaker button next to the field is pressed. The button sends the phrase to the provider of the selected account, with the LLM (model), Voice and Extra instruction chosen above, and plays the returned audio.
- If the request fails, the error reported by the provider is shown in this field instead of the phrase. Typical causes are a wrong API key, a URL that does not match the selected API, or a voice the provider does not know.
Add an app
- TextToSpeech API - the app object other apps use to reach the speech service
- Name
- The name displayed for the app object which must be unique.
- SIP
- The sip from the app object which must be unique.
Existing configuration templates will be listed, app objects can be assigned by checking the respective checkbox.
Be Aware: Every user who works with an app that produces speech output needs the TextToSpeech API assigned as well (by using a config template, for example), because that app reaches the speech service through the user's own app list.