# voice-chat **Repository Path**: kpret/voice-chat ## Basic Information - **Project Name**: voice-chat - **Description**: No description available - **Primary Language**: Unknown - **License**: Not specified - **Default Branch**: main - **Homepage**: None - **GVP Project**: No ## Statistics - **Stars**: 0 - **Forks**: 0 - **Created**: 2026-04-02 - **Last Updated**: 2026-04-06 ## Categories & Tags **Categories**: Uncategorized **Tags**: None ## README # Voice Chat Assistant A Java 21 + Spring Boot voice assistant with an Apple-inspired browser UI. The app records audio in the browser, sends it to a Java backend, runs STT → LLM → TTS, then echoes the transcript and assistant reply back onto the page with reply audio playback. ## Features - Spring Boot + Maven backend - Spring AI + Spring AI OpenAI integration - Apple-style glassmorphism UI - Browser microphone capture, transcript echo, and reply playback - OpenAI-backed STT / chat / TTS integration - Xiaomi MiMo `mimo-v2-tts` integration for speech synthesis - Mock mode for local development without cloud credentials - Separate base URL / API key settings for chat, transcription, and speech - Verified backend tests and packaging ## Project layout ```text src/main/java/com/voicechat/assistant/ application, controllers, services, websocket src/main/resources/static/ Apple-style frontend assets src/test/java/com/voicechat/assistant/ controller + service tests .env.example local configuration template docs/integration-review.md verification notes ``` ## Configuration 1. Copy the template: ```bash cp .env.example .env ``` 2. For local demo mode, keep: ```env APP_VOICE_PROVIDER_MOCK=true ``` 3. For a live provider, configure each model type separately: ```env APP_VOICE_PROVIDER_MOCK=false APP_VOICE_CHAT_BASE_URL=https://api.openai.com APP_VOICE_CHAT_API_KEY=your_chat_api_key APP_VOICE_TRANSCRIPTION_BASE_URL=https://api.openai.com APP_VOICE_TRANSCRIPTION_API_KEY=your_transcription_api_key APP_VOICE_SPEECH_BASE_URL=https://api.openai.com APP_VOICE_SPEECH_API_KEY=your_speech_api_key ``` 4. To use Xiaomi MiMo for TTS, keep chat / STT as-is and only switch the speech settings: ```env APP_VOICE_PROVIDER_MOCK=false APP_VOICE_SPEECH_BASE_URL=https://api.xiaomimimo.com APP_VOICE_SPEECH_API_KEY=your_mimo_api_key APP_VOICE_SPEECH_MODEL=mimo-v2-tts APP_VOICE_SPEECH_VOICE=mimo_default APP_VOICE_SPEECH_FORMAT=wav ``` 5. Keep `.env` local only; it is ignored by git. ### Framework note The backend now uses: - `Spring AI` for the chat / transcription / speech abstractions - `Spring AI OpenAI` for the OpenAI model implementations The code intentionally keeps the three model connections separate so different model types can use different credentials or endpoints. For Xiaomi MiMo TTS, the backend automatically switches from the OpenAI `/audio/speech` flow to Xiaomi's OpenAI-compatible `POST /v1/chat/completions` audio flow when `APP_VOICE_SPEECH_MODEL=mimo-v2-tts`. ## Run ```bash mvn spring-boot:run ``` Open: - http://localhost:8080 ## HTTP flow - `POST /api/voice/chat` — browser-friendly multipart endpoint used by the UI - `GET /api/v1/voice/health` — backend health - `GET /api/v1/voice/config` — backend config summary - `POST /api/v1/voice/transcriptions` — direct STT endpoint - `POST /api/v1/voice/turns` — transcript-to-reply endpoint - `WS /ws/voice` — text-turn websocket contract ## Verification ```bash mvn test mvn -DskipTests package ``` Worker verification also reported mock curl checks for `/health`, `/config`, and `/turns`. See `docs/integration-review.md` for the consolidated notes. ## Current limitations - The browser UI is turn-based rather than full streaming audio. - End-to-end live provider behavior still depends on valid provider credentials in `.env`. - Browser microphone permission and autoplay behavior vary by browser.