📄️AudioClient.PlayStream() will reject audio that looks fine but isn'tSymptomPitfall📄️ElevenLabs viseme timestamps are per-chunk, not per-utterance — and need a sample-rate correction on topSymptomPitfall📄️The dedicated kill-phrase mic plan died at the USB port countThe situationPitfall📄️Mic gain is two independent settings, not oneSymptomPitfall📄️openWakeWord's own startup noise can silently wreck your false-positive testingSymptomPitfall📄️Pipecat's SegmentedSTTService silently double-wraps audio unless you override one flagSymptomPitfall📄️STT upgrade paths under evaluation — nothing decided yetThis is an honest snapshot of options actually researched, not a plan that's been started. Nothing below has been built or benchmarked on this project's actual hardware yet. Filed here as an open question, not a decision.📄️Reusing the cloud-primary/local-fallback pattern for TTS, not just LLMThe situationDecision📄️Amplitude-based lip sync over phoneme-based — a scope decision, not a quality oneThe situationDecision
📄️ElevenLabs viseme timestamps are per-chunk, not per-utterance — and need a sample-rate correction on topSymptomPitfall
📄️Pipecat's SegmentedSTTService silently double-wraps audio unless you override one flagSymptomPitfall
📄️STT upgrade paths under evaluation — nothing decided yetThis is an honest snapshot of options actually researched, not a plan that's been started. Nothing below has been built or benchmarked on this project's actual hardware yet. Filed here as an open question, not a decision.
📄️Amplitude-based lip sync over phoneme-based — a scope decision, not a quality oneThe situationDecision