feat(voice): configure independent speech endpoints - #359
Conversation
morgmart
left a comment
There was a problem hiding this comment.
🤖 Automated code review
REQUEST_CHANGES: two blocking regressions remain in the endpoint migration and reset lifecycle. Supplied GitHub evidence shows all ten captured checks completed successfully, but required checks still govern merge readiness.
Deterministic publication result: 2 blocking and 0 non-blocking inline finding(s) publishable; 0 duplicate(s) suppressed; 0 blocking screenshot-evidence requirement(s) in this review body.
morgmart
left a comment
There was a problem hiding this comment.
🤖 Automated code review
APPROVE: the exact PR comparison has no publishable Engineering findings. The two prior blocking issues are fixed in the current head and their resolved threads contain substantive author replies. Supplied GitHub evidence is structurally valid; six captured checks passed and two were still in progress, so required checks continue to govern merge readiness.
Deterministic publication result: 0 blocking and 0 non-blocking inline finding(s) publishable; 0 duplicate(s) suppressed; 0 blocking screenshot-evidence requirement(s) in this review body.
Pending checks: 2 check(s) are not complete.
This approval reflects the completed code review only; merge readiness remains governed by the repository's required checks.
morgmart
left a comment
There was a problem hiding this comment.
🤖 Automated code review
APPROVE: the exact PR comparison has no publishable Engineering findings. The two prior blocking issues remain fixed, and the latest endpoint-selection change correctly preserves an explicit default OpenAI URL as an override of legacy environment routing with discriminating coverage. All ten captured GitHub checks completed successfully; required checks still govern merge readiness.
Deterministic publication result: 0 blocking and 0 non-blocking inline finding(s) publishable; 0 duplicate(s) suppressed; 0 blocking screenshot-evidence requirement(s) in this review body.
Pending checks: 1 check(s) are not complete.
This approval reflects the completed code review only; merge readiness remains governed by the repository's required checks.
Summary
Teams using local or OpenAI-compatible speech services need to route Realtime, transcription, and synthesis independently without replacing the key used by the default OpenAI endpoints.
Adds independent full URL settings for OpenAI-compatible Realtime, speech-to-text, and text-to-speech endpoints. Each service has one Save action for its URL and key. Keys are stored per URL in macOS Keychain, while the default OpenAI URLs share one key. Settings show “No key set” or a dotted saved-key placeholder without reading the secret. Blank URLs use the OpenAI defaults. Existing STT and TTS routing through
BERD_OPENAI_VOICE_BASE_URLretains its shared credential; explicitly saved URLs use their own URL-scoped keys. Native TTS, STT, and Realtime startup capture the destination and its URL-scoped credential together, so a concurrent endpoint save cannot send a key to a different endpoint.berd-callalso accepts session-only--realtime-url,--stt-url, and--tts-urloverrides. Realtime dictation continues to use OpenAI's fixed client-secret API; a custom Realtime URL applies to Expert-Spokesperson conversations.Related issue
N/A
Testing
ws://127.0.0.1:18870/v1/realtime?intent=transcriptionfor STT andhttp://127.0.0.1:18870/v1/audio/speechfor TTS. Enter a disposable key such aslocal-testin each key field and click that service's single Save button. No server is needed to verify settings.ws://127.0.0.1:18870/v1/realtime. It has its own saved-key status. Clear the disposable keys with Remove, then blank each custom URL and Save to return to the OpenAI defaults.Endpoint routing and call-scoped overrides
After building
berd-call, run from the checkout:The script starts local STT, TTS, and Realtime fixtures and runs the CLI against each override URL with a disposable key and a PCM test host. It then starts a new call without URL overrides. Observed requests are
/stt,/tts,/realtime, then/default-sttand/default/audio/speech. Synthesis completes, the saved endpoint settings remain unchanged, and the new call uses its defaults. No microphone or speakers are opened.Screenshots
STT and TTS before: shared OpenAI key
STT and TTS after: independent endpoint controls
Voice assistant before: shared OpenAI key
Voice assistant after: custom Realtime endpoint and key