Voice Library

Search, compare, and choose provider-backed voices in the OpenPhonex workspace.

The Voice Library is the workspace catalog for voices that OpenPhonex has already recorded from supported provider inventory. Use it to compare a voice's provider, models, declared languages, accent and style metadata before choosing one for a hosted agent.

It is a browse surface, not a provider-control surface: opening the library does not contact a provider, refresh inventory, retrieve credentials, or change an agent.

Search the catalog

Search and filters apply to the complete cached catalog, not only the cards already visible on the current page. You can combine these filters:

  • Language
  • Provider
  • Accent
  • Character or style
  • Gender
  • Free-text search

Selecting several values in one filter means “any of these values.” Selecting values in different filters means “must satisfy each filter.” For example, selecting Estonian and Finnish in Language, plus two providers, returns a voice from either language and either provider; adding a calm Character requires a matching style as well.

Voice language claims remain model-specific. When a provider gives no per-voice language list, OpenPhonex uses only that model's declared languages; it does not infer a language from another voice with the same provider or name.

Pages and model groups

The library opens with up to 24 voice groups per page and accepts a page size up to 100. When the current cached result count is available, the pager shows Page X of Y, numbered links, and Next and Previous. Filters and page position stay in the URL so you can return to the same view. Changing a filter starts at the first matching page; an old bookmarked page that is now beyond the result set returns to the first page instead of leaving you on a blank one.

The page count is the number of currently browsable, filtered voice groups in the cache. It is not a count of models and it is not a claim about a provider's complete or undiscovered inventory. When that count is not available during a rolling update, the page keeps the compatible Next and Previous controls instead of inventing a total.

Models are grouped before pagination only when the provider's inventory evidence identifies them as the same voice. A matching raw identifier alone is not enough: two models, or two providers, can expose distinct voices with the same name or ID. This keeps a model-specific selection exact while avoiding duplicate cards for a provider-evidenced shared voice.

If an agent already uses a voice that is outside the editor's small embedded gallery, the editor resolves that exact stored provider, model and voice ID instead of treating it as missing just because it is on a later library page.

What catalog coverage means

The library keeps the last successful cached inventory visible while a refresh is running or has failed. A partial or ranked provider search can show useful results without proving that it found every voice the provider offers. When the workspace says the accessible provider window is limited, treat the library as the available cached window—not as a claim that no other provider voices exist.

Only a completed authoritative provider refresh can remove a previously cached voice for absence. Provider eligibility, licensing and your agent's existing selection continue to be checked when you save or run the agent; appearing in a search result is not permission to bypass those controls.

Preview a voice in the intended language

Use Play to audition a voice without changing an agent or refreshing a provider catalog. Browsing and changing filters do not create a preview or use your preview allowance.

If one active Language filter matches a voice, Play targets that language. When several selected languages match the same voice, the row shows a searchable, single-language chooser with the provider's declared values. Choose one before a live preview starts; OpenPhonex never silently takes the first matching language. With no Language filter, an existing recorded sample can still play as published until you explicitly choose a preview language. If its catalog provenance does not establish a spoken language, the row labels that sample as unverified rather than inferring one.

When a target language is selected, a recorded sample is used only when its catalog provenance proves that it matches the target. At present, an OpenPhonex ingestion-synthesized sample has that proof for English; a provider-supplied recording has no spoken-language assertion and is not treated as English or as a match for another language. The row uses the bounded live preview instead.

A preview can also carry a delivery profile (natural, lively, calm or precise) so you hear the voice performed the way a call with that profile would perform it; the response says whether the selected provider and model applied it.

Text you type is sent and spoken exactly as written by text-to-speech voices (GPT-Live voices use a fixed sample sentence instead; see below). OpenPhonex does not translate it. If you leave the text blank, the library uses a localized sample sentence for the selected language where one is available. If there is no such sentence, enter your own text rather than receiving an English fallback. Live preview remains bounded, metered, and rate-limited for the workspace.

OpenAI GPT-Live 1 voices

The fourteen GPT-Live 1 voices are listed under the OpenAI provider. They are conversation-engine voices, not text-to-speech voices, and a few things differ:

  • Every language filter shows them, in a Multilingual group. The engine follows the caller's language automatically, so no language filter can exclude a GPT-Live voice and no per-voice language list is shown. The group says: "Follows the caller's language automatically. Widely supported languages are recognised reliably. Some languages may be misheard and answered in another language, so test yours before going live." When the library is filtered to Sinhala, one extra note asks you to test first. The dated results behind both sentences are on Which languages the voice agent follows.
  • Details are OpenAI's words. Gender comes from OpenAI's presentation (Feminine or Masculine). The row's details show OpenAI's regional influence (which OpenAI describes as speaking style, not accent fidelity), its per-voice Language label (English or Portuguese, a description of the voice, not a limit on what it speaks) and whether the voice is Natural or Generated. Marin and Cedar are OpenAI's undocumented voices: they show a presentation and nothing else.
  • Play runs a short real GPT-Live session. The clip is rendered once per voice and sample sentence and cached, so the same audition plays instantly next time. It uses a fixed OpenPhonex sample sentence for the chosen preview language; text you type is not sent to a GPT-Live voice, and the engine paraphrases rather than reading verbatim. Previews are not charged to your workspace and never count as call usage.
  • Both figures say what they measure. The latency is the time to first audio on a scripted greeting, measured by OpenPhonex on the date the row shows; reply latency after a caller speaks is not measured and is not shown as zero. The price is call duration at the per-minute rate, plus the backend model's tokens billed separately per call; see Conversation engines for the rate.
  • Use switches the engine. Choosing a GPT-Live voice for an agent sets the agent's conversation engine to GPT-Live 1 and says so. The change is a draft until you publish, exactly as any other voice change.

Choose a voice for an agent

Choose a provider, model and voice together. The agent configuration flow validates that exact combination for the agent's language and credential scope before it saves. Bring-your-own-key voices can be private to your own provider account and may not appear in the shared library; OpenPhonex keeps the stored selection distinct from shared public inventory rather than exposing a private provider listing to other workspaces.

On this page