{"broker_url":"http://192.168.1.33:8080","hosts":[{"host":"ai-150","agent_alive":false,"cpu":"AMD Ryzen 7 9800X3D","ram_mib_total":98304,"ram_mib_budget":80000,"ram_mib_used":13410,"gpus":[{"index":0,"name":"NVIDIA GeForce RTX 5090","vram_mib_budget":30000,"mem_total_mib":32607,"mem_used_mib":2052,"util_pct":6},{"index":1,"name":"NVIDIA GeForce RTX 3090","vram_mib_budget":22000,"mem_total_mib":24576,"mem_used_mib":798,"util_pct":7}],"variants":[{"service":"qwen3-30b-a3b-q4__gpu0","fidelity":null,"capabilities":["qwen3-30b-a3b-q4"],"placement":"gpu0","gpus":[0],"vram_mib":{"0":23000},"cpu_mib":0,"exclusive_with":[],"model":{"name":"Qwen3-30B-A3B-Instruct-2507","params":"30B (MoE, 3B active)","quant":"Q4_K_M","context":32768,"vision":false,"source":"https://huggingface.co/unsloth/Qwen3-30B-A3B-Instruct-2507-GGUF"}},{"service":"mistral-small-3.2-24b-q4__gpu1","fidelity":null,"capabilities":["mistral-small-3.2-24b-q4"],"placement":"gpu1","gpus":[1],"vram_mib":{"1":19000},"cpu_mib":0,"exclusive_with":[],"model":{"name":"Mistral-Small-3.2-24B-Instruct-2506","params":"24B","quant":"Q4_K_M","context":32768,"vision":false,"source":"https://huggingface.co/unsloth/Mistral-Small-3.2-24B-Instruct-2506-GGUF"}},{"service":"qwen3-vl-32b-q8__all","fidelity":null,"capabilities":["qwen3-vl-32b-q8"],"placement":"all","gpus":[0,1],"vram_mib":{"0":28500,"1":18500},"cpu_mib":0,"exclusive_with":[],"model":{"name":"Qwen3-VL-32B-Instruct","params":"32B (dense)","quant":"Q8_0","context":49152,"vision":true,"source":"https://huggingface.co/unsloth/Qwen3-VL-32B-Instruct-GGUF","notes":"Offline deep-work tier. Spans both GPUs (tensor-split 60/40). Q8 weights + 48K ctx q8 KV. Higher capability than the live tier; evicted instantly when any live capability is requested."}},{"service":"qwen3-vl-32b-q4__gpu0","fidelity":null,"capabilities":["qwen3-vl-32b-q4"],"placement":"gpu0","gpus":[0],"vram_mib":{"0":28000},"cpu_mib":0,"exclusive_with":[],"model":{"name":"Qwen3-VL-32B-Instruct","params":"32B","quant":"Q4_K_M","context":32768,"vision":true,"source":"https://huggingface.co/unsloth/Qwen3-VL-32B-Instruct-GGUF","notes":"Provisional vision-fast lane. Same model as deep_split but Q4 and gpu0-only, so it evicts the live director when used. Current use case (scene-image verify/caption) is bursty enough that this is fine; the eviction cost slots in between live calls. When portrait localization returns or extraction runs at scale, replace this with a genuinely small VLM (Qwen2-VL-2B / Moondream2 / SmolVLM2) on the .151 3090 so high-frequency cheap reads never touch the live hot path. Per the brief: vision-doc = low-frequency / high-stakes / fidelity-critical (32B is right); vision-fast = high-frequency / low-stakes / latency-sensitive (small + isolated lane is right)."}},{"service":"comfyui-engine__gpu0","fidelity":null,"capabilities":["comfyui-engine","comfyui-flux-unchained","comfyui-map-sdxl","comfyui-sprite-rgba"],"placement":"gpu0","gpus":[0],"vram_mib":{"0":16000},"cpu_mib":0,"exclusive_with":[],"model":{"name":"ComfyUI","notes":"Workflow-driven; loads checkpoints at runtime. flux1-dev-Q5_K_S + t5-xxl present in ComfyUI/models/. Burst tier: scene imagery queues until prose drains, never displaces live."}},{"service":"comfyui-engine__gpu1","fidelity":null,"capabilities":["comfyui-engine","comfyui-flux-unchained","comfyui-map-sdxl","comfyui-sprite-rgba"],"placement":"gpu1","gpus":[1],"vram_mib":{"1":16000},"cpu_mib":0,"exclusive_with":[],"model":{"name":"ComfyUI","notes":"Workflow-driven; loads checkpoints at runtime. flux1-dev-Q5_K_S + t5-xxl present in ComfyUI/models/. Burst tier: scene imagery queues until prose drains, never displaces live."}},{"service":"qwen3-embedding-4b-q8__cpu","fidelity":null,"capabilities":["qwen3-embedding-4b-q8"],"placement":"cpu","gpus":[],"vram_mib":{},"cpu_mib":10000,"exclusive_with":[],"model":{"name":"Qwen3-Embedding-4B","params":"4B","quant":"Q8_0","context":32768,"vision":false,"source":"https://huggingface.co/Qwen/Qwen3-Embedding-4B-GGUF"}},{"service":"bge-reranker-v2-m3-q8__cpu","fidelity":null,"capabilities":["bge-reranker-v2-m3-q8"],"placement":"cpu","gpus":[],"vram_mib":{},"cpu_mib":1500,"exclusive_with":[],"model":{"name":"BGE-reranker-v2-m3","params":"568M","quant":"Q8_0","context":8192,"vision":false,"source":"https://huggingface.co/gpustack/bge-reranker-v2-m3-GGUF"}}]},{"host":"ai-151","agent_alive":false,"cpu":"AMD Ryzen 7 3800X","ram_mib_total":32689,"ram_mib_budget":24000,"ram_mib_used":10251,"gpus":[{"index":0,"name":"NVIDIA GeForce RTX 3090","vram_mib_budget":22000,"mem_total_mib":24576,"mem_used_mib":1650,"util_pct":19},{"index":1,"name":"NVIDIA GeForce RTX 2080 Ti","vram_mib_budget":10000,"mem_total_mib":11264,"mem_used_mib":0,"util_pct":0}],"variants":[{"service":"qwen3-30b-a3b-q4__gpu0","fidelity":null,"capabilities":["qwen3-30b-a3b-q4"],"placement":"gpu0","gpus":[0],"vram_mib":{"0":21000},"cpu_mib":0,"exclusive_with":[],"model":{"name":"Qwen3-30B-A3B-Instruct-2507","params":"30B (MoE, 3B active)","quant":"Q4_K_M","context":32768,"vision":false,"source":"https://huggingface.co/unsloth/Qwen3-30B-A3B-Instruct-2507-GGUF"}},{"service":"mistral-small-3.2-24b-q4__gpu0","fidelity":null,"capabilities":["mistral-small-3.2-24b-q4"],"placement":"gpu0","gpus":[0],"vram_mib":{"0":19000},"cpu_mib":0,"exclusive_with":[],"model":{"name":"Mistral-Small-3.2-24B-Instruct-2506","params":"24B","quant":"Q4_K_M","context":32768,"vision":false,"source":"https://huggingface.co/unsloth/Mistral-Small-3.2-24B-Instruct-2506-GGUF"}},{"service":"qwen3-vl-32b-q4__gpu0","fidelity":null,"capabilities":["qwen3-vl-32b-q4"],"placement":"gpu0","gpus":[0],"vram_mib":{"0":22000},"cpu_mib":0,"exclusive_with":[],"model":{"name":"Qwen3-VL-32B-Instruct","params":"32B","quant":"Q4_K_M","context":32768,"vision":true,"source":"https://huggingface.co/unsloth/Qwen3-VL-32B-Instruct-GGUF","notes":"Provisional vision-fast lane. Same model as deep_split but Q4 and gpu0-only, so it evicts the live director when used. Current use case (scene-image verify/caption) is bursty enough that this is fine; the eviction cost slots in between live calls. When portrait localization returns or extraction runs at scale, replace this with a genuinely small VLM (Qwen2-VL-2B / Moondream2 / SmolVLM2) on the .151 3090 so high-frequency cheap reads never touch the live hot path. Per the brief: vision-doc = low-frequency / high-stakes / fidelity-critical (32B is right); vision-fast = high-frequency / low-stakes / latency-sensitive (small + isolated lane is right)."}},{"service":"comfyui-engine__gpu0","fidelity":null,"capabilities":["comfyui-engine","comfyui-flux-unchained","comfyui-map-sdxl","comfyui-sprite-rgba"],"placement":"gpu0","gpus":[0],"vram_mib":{"0":16000},"cpu_mib":0,"exclusive_with":[],"model":{"name":"ComfyUI","notes":"Workflow-driven; loads checkpoints at runtime. flux1-dev-Q5_K_S + t5-xxl present in ComfyUI/models/. Burst tier: scene imagery queues until prose drains, never displaces live."}},{"service":"qwen3-embedding-4b-q8__gpu0","fidelity":null,"capabilities":["qwen3-embedding-4b-q8"],"placement":"gpu0","gpus":[0],"vram_mib":{"0":5500},"cpu_mib":0,"exclusive_with":[],"model":{"name":"Qwen3-Embedding-4B","params":"4B","quant":"Q8_0","context":32768,"vision":false,"source":"https://huggingface.co/Qwen/Qwen3-Embedding-4B-GGUF"}},{"service":"faster-whisper__gpu1","fidelity":null,"capabilities":["faster-whisper"],"placement":"gpu1","gpus":[1],"vram_mib":{"1":4000},"cpu_mib":0,"exclusive_with":[],"model":{"name":"faster-whisper @ Zoont/faster-whisper-large-v3-turbo-int8-ct2","notes":"Wrapped by stt_service (uvicorn app:app on the shared embedded python). Clients use `switchboard_client.transcribe(audio, ...)` — it handles ticket acquisition, multipart upload, and form fields. POST /transcribe accepts multipart `audio` + form fields (initial_prompt, language, model, word_timestamps); returns text + per-segment + per-word timing + no_speech/low_confidence flags. Default model swapped (Jun 5) from `large-v3-turbo` (runtime-quantized to int8_float16) to `Zoont/faster-whisper-large-v3-turbo-int8-ct2` (pre-quantized int8 weights from HuggingFace). Same Whisper model under the hood; pre-quantized weights load slightly faster and use less VRAM. Compute kept at int8_float16 for GPU mixed-precision inference. Default upload cap 10 MB (STT_MAX_UPLOAD_BYTES). Verified Jun 5: same transcription quality + sub-second inference on 2-6 s clips. STREAMING /transcribe_stream WebSocket added Jun 5 for real-time turn-taking (sb.transcribe_stream in client v1.2.0+). Server-side Silero VAD endpointing; partial+final events. Jun 5 fix: hotwords now passed via faster-whisper's native `hotwords=` (CTranslate2 logit biasing) instead of concatenated into initial_prompt -- the latter caused Whisper to echo the vocab list as the first partial of every utterance. Two backstops added: drop partials with no_speech_prob>threshold (cough/breath suppression), and _is_prompt_echo() check that suppresses any output text equal to the hotwords/prompt verbatim. Jun 5 followup: native hotwords still hallucinated the vocab list on unclear mic audio (DM AI). REPLACED the output-filter approach with TWO-PASS DECODE ARBITRATION in _decode(): when hotwords are configured, runs whisper twice (baseline=no-bias, biased=with hotwords) and arbitrates. Baseline empty/no_speech -> use baseline. Biased introduced hotwords baseline never saw -> use baseline (bias was artificial). Otherwise -> use biased (baseline confirms hotwords belong, biased just spells them better). Cost: 2x decode time per partial when hotwords configured. Source-level fix; legitimate one-word hotword answers (\"Karrthûn\" alone) pass through."}},{"service":"kokoro__gpu1","fidelity":null,"capabilities":["kokoro"],"placement":"gpu1","gpus":[1],"vram_mib":{"1":1500},"cpu_mib":0,"exclusive_with":[],"model":{"name":"Kokoro-82M (hexgrad/Kokoro-82M, Apache-2.0)","notes":"Canonical TTS for the fleet (replaced AllTalk on cutover Jun 5). 82M-param StyleTTS2 distillation; ~190ms warm render for a 6.5s line (~33x realtime). 24kHz mono float32 output. Native `speed` parameter (0.5-2.0). 50+ voicepacks via /voices. Sidecar at C:\\AI\\kokoro_service installs the kokoro pip package into the existing C:\\AI\\python-3.12.10-embed-amd64 distribution (shared with agent + stt + diarization; torch 2.5.1+cu121 was already there). transformers pinned to <5.0 (5.x requires torch>=2.7). LANGUAGE EXTRAS: `kokoro` pulls `misaki[en]` only; non-English pipelines lazy-fail on import unless their misaki extras are installed too. install.ps1 explicitly pip installs `misaki[ja]`, `misaki[zh]`, `misaki[ko]`, and `ordered_set` to cover all currently-exposed non-English voicepacks (the Mandarin `lang=z` pipeline in particular imports `ordered_set` via `misaki.zh.transcription`). KPipeline init passes `repo_id='hexgrad/Kokoro-82M'` explicitly to suppress the noisy default-repo warning misaki logs per fresh language load. skip_if_missing keys on a .installed marker the install script touches at the end. Clients should use switchboard_client.tts_say* — direct /generate POSTs are supported but bypass the chunking and tag system."}},{"service":"pyannote__gpu1","fidelity":null,"capabilities":["pyannote"],"placement":"gpu1","gpus":[1],"vram_mib":{"1":2000},"cpu_mib":0,"exclusive_with":[],"model":{"name":"pyannote-audio speaker-diarization-3.1","notes":"Local model files only (HF_HUB_OFFLINE)."}},{"service":"gemma-4-26b-a4b-q4__gpu0","fidelity":null,"capabilities":["gemma-4-26b-a4b-q4"],"placement":"gpu0","gpus":[0],"vram_mib":{"0":17000},"cpu_mib":0,"exclusive_with":[],"model":{"name":"gemma-4-26B-A4B-it","params":"26B (MoE, 4B active)","quant":"Q4_K_M","context":16384,"notes":"Wired from a previously-orphaned copy on ai-151. Offline tier (swap-in)."}}]},{"host":"ai-9","agent_alive":true,"cpu":"Intel Core i7-4790K","ram_mib_total":32689,"ram_mib_budget":24000,"ram_mib_used":8072,"gpus":[{"index":0,"name":"NVIDIA GeForce GTX 1080 Ti","vram_mib_budget":10000,"mem_total_mib":11264,"mem_used_mib":736,"util_pct":0}],"variants":[{"service":"qwen3-8b-q4__gpu0","fidelity":null,"capabilities":["qwen3-8b-q4"],"placement":"gpu0","gpus":[0],"vram_mib":{"0":8500},"cpu_mib":0,"exclusive_with":[],"model":{"name":"Qwen3-8B (Qwen/Qwen3-8B-GGUF, Q4_K_M)","notes":"Small-LLM lane on the 1080 Ti via llama.cpp VULKAN backend (fleet's CUDA build excludes Pascal sm_61). 32K ctx + q8_0 KV ~8.1GiB.","context":32768}}]}],"artifacts":[{"id":"qwen3-30b-a3b-q4","kind":"llamacpp","model":{"name":"Qwen3-30B-A3B-Instruct-2507","params":"30B (MoE, 3B active)","quant":"Q4_K_M","context":32768,"vision":false,"source":"https://huggingface.co/unsloth/Qwen3-30B-A3B-Instruct-2507-GGUF"},"engine_ref":null,"profiles":[{"host":"ai-150","gpus":[0],"placement":"gpu0","min_warm_sec":60,"max_context":32768},{"host":"ai-151","gpus":[0],"placement":"gpu0","min_warm_sec":60,"max_context":16384}]},{"id":"mistral-small-3.2-24b-q4","kind":"llamacpp","model":{"name":"Mistral-Small-3.2-24B-Instruct-2506","params":"24B","quant":"Q4_K_M","context":32768,"vision":false,"source":"https://huggingface.co/unsloth/Mistral-Small-3.2-24B-Instruct-2506-GGUF"},"engine_ref":null,"profiles":[{"host":"ai-150","gpus":[1],"placement":"gpu1","min_warm_sec":60,"max_context":32768},{"host":"ai-151","gpus":[0],"placement":"gpu0","min_warm_sec":60,"max_context":32768}]},{"id":"qwen3-vl-32b-q8","kind":"llamacpp","model":{"name":"Qwen3-VL-32B-Instruct","params":"32B (dense)","quant":"Q8_0","context":49152,"vision":true,"source":"https://huggingface.co/unsloth/Qwen3-VL-32B-Instruct-GGUF","notes":"Offline deep-work tier. Spans both GPUs (tensor-split 60/40). Q8 weights + 48K ctx q8 KV. Higher capability than the live tier; evicted instantly when any live capability is requested."},"engine_ref":null,"profiles":[{"host":"ai-150","gpus":[0,1],"placement":"all","min_warm_sec":120,"max_context":49152}]},{"id":"qwen3-vl-32b-q4","kind":"llamacpp","model":{"name":"Qwen3-VL-32B-Instruct","params":"32B","quant":"Q4_K_M","context":32768,"vision":true,"source":"https://huggingface.co/unsloth/Qwen3-VL-32B-Instruct-GGUF","notes":"Provisional vision-fast lane. Same model as deep_split but Q4 and gpu0-only, so it evicts the live director when used. Current use case (scene-image verify/caption) is bursty enough that this is fine; the eviction cost slots in between live calls. When portrait localization returns or extraction runs at scale, replace this with a genuinely small VLM (Qwen2-VL-2B / Moondream2 / SmolVLM2) on the .151 3090 so high-frequency cheap reads never touch the live hot path. Per the brief: vision-doc = low-frequency / high-stakes / fidelity-critical (32B is right); vision-fast = high-frequency / low-stakes / latency-sensitive (small + isolated lane is right)."},"engine_ref":null,"profiles":[{"host":"ai-150","gpus":[0],"placement":"gpu0","min_warm_sec":90,"max_context":32768},{"host":"ai-151","gpus":[0],"placement":"gpu0","min_warm_sec":90,"max_context":8192}]},{"id":"comfyui-engine","kind":"comfyui-engine","model":{"name":"ComfyUI","notes":"Workflow-driven; loads checkpoints at runtime. flux1-dev-Q5_K_S + t5-xxl present in ComfyUI/models/. Burst tier: scene imagery queues until prose drains, never displaces live."},"engine_ref":null,"profiles":[{"host":"ai-150","gpus":[0],"placement":"gpu0","min_warm_sec":30,"max_context":null},{"host":"ai-150","gpus":[1],"placement":"gpu1","min_warm_sec":30,"max_context":null},{"host":"ai-151","gpus":[0],"placement":"gpu0","min_warm_sec":30,"max_context":null}]},{"id":"qwen3-embedding-4b-q8","kind":"embed","model":{"name":"Qwen3-Embedding-4B","params":"4B","quant":"Q8_0","context":32768,"vision":false,"source":"https://huggingface.co/Qwen/Qwen3-Embedding-4B-GGUF"},"engine_ref":null,"profiles":[{"host":"ai-150","gpus":[],"placement":"cpu","min_warm_sec":20,"max_context":32768},{"host":"ai-151","gpus":[0],"placement":"gpu0","min_warm_sec":20,"max_context":32768}]},{"id":"bge-reranker-v2-m3-q8","kind":"rerank","model":{"name":"BGE-reranker-v2-m3","params":"568M","quant":"Q8_0","context":8192,"vision":false,"source":"https://huggingface.co/gpustack/bge-reranker-v2-m3-GGUF"},"engine_ref":null,"profiles":[{"host":"ai-150","gpus":[],"placement":"cpu","min_warm_sec":15,"max_context":8192}]},{"id":"faster-whisper","kind":"stt","model":{"name":"faster-whisper @ Zoont/faster-whisper-large-v3-turbo-int8-ct2","notes":"Wrapped by stt_service (uvicorn app:app on the shared embedded python). Clients use `switchboard_client.transcribe(audio, ...)` — it handles ticket acquisition, multipart upload, and form fields. POST /transcribe accepts multipart `audio` + form fields (initial_prompt, language, model, word_timestamps); returns text + per-segment + per-word timing + no_speech/low_confidence flags. Default model swapped (Jun 5) from `large-v3-turbo` (runtime-quantized to int8_float16) to `Zoont/faster-whisper-large-v3-turbo-int8-ct2` (pre-quantized int8 weights from HuggingFace). Same Whisper model under the hood; pre-quantized weights load slightly faster and use less VRAM. Compute kept at int8_float16 for GPU mixed-precision inference. Default upload cap 10 MB (STT_MAX_UPLOAD_BYTES). Verified Jun 5: same transcription quality + sub-second inference on 2-6 s clips. STREAMING /transcribe_stream WebSocket added Jun 5 for real-time turn-taking (sb.transcribe_stream in client v1.2.0+). Server-side Silero VAD endpointing; partial+final events. Jun 5 fix: hotwords now passed via faster-whisper's native `hotwords=` (CTranslate2 logit biasing) instead of concatenated into initial_prompt -- the latter caused Whisper to echo the vocab list as the first partial of every utterance. Two backstops added: drop partials with no_speech_prob>threshold (cough/breath suppression), and _is_prompt_echo() check that suppresses any output text equal to the hotwords/prompt verbatim. Jun 5 followup: native hotwords still hallucinated the vocab list on unclear mic audio (DM AI). REPLACED the output-filter approach with TWO-PASS DECODE ARBITRATION in _decode(): when hotwords are configured, runs whisper twice (baseline=no-bias, biased=with hotwords) and arbitrates. Baseline empty/no_speech -> use baseline. Biased introduced hotwords baseline never saw -> use baseline (bias was artificial). Otherwise -> use biased (baseline confirms hotwords belong, biased just spells them better). Cost: 2x decode time per partial when hotwords configured. Source-level fix; legitimate one-word hotword answers (\"Karrthûn\" alone) pass through."},"engine_ref":null,"profiles":[{"host":"ai-151","gpus":[1],"placement":"gpu1","min_warm_sec":15,"max_context":null}]},{"id":"kokoro","kind":"tts","model":{"name":"Kokoro-82M (hexgrad/Kokoro-82M, Apache-2.0)","notes":"Canonical TTS for the fleet (replaced AllTalk on cutover Jun 5). 82M-param StyleTTS2 distillation; ~190ms warm render for a 6.5s line (~33x realtime). 24kHz mono float32 output. Native `speed` parameter (0.5-2.0). 50+ voicepacks via /voices. Sidecar at C:\\AI\\kokoro_service installs the kokoro pip package into the existing C:\\AI\\python-3.12.10-embed-amd64 distribution (shared with agent + stt + diarization; torch 2.5.1+cu121 was already there). transformers pinned to <5.0 (5.x requires torch>=2.7). LANGUAGE EXTRAS: `kokoro` pulls `misaki[en]` only; non-English pipelines lazy-fail on import unless their misaki extras are installed too. install.ps1 explicitly pip installs `misaki[ja]`, `misaki[zh]`, `misaki[ko]`, and `ordered_set` to cover all currently-exposed non-English voicepacks (the Mandarin `lang=z` pipeline in particular imports `ordered_set` via `misaki.zh.transcription`). KPipeline init passes `repo_id='hexgrad/Kokoro-82M'` explicitly to suppress the noisy default-repo warning misaki logs per fresh language load. skip_if_missing keys on a .installed marker the install script touches at the end. Clients should use switchboard_client.tts_say* — direct /generate POSTs are supported but bypass the chunking and tag system."},"engine_ref":null,"profiles":[{"host":"ai-151","gpus":[1],"placement":"gpu1","min_warm_sec":15,"max_context":null}]},{"id":"pyannote","kind":"diarization","model":{"name":"pyannote-audio speaker-diarization-3.1","notes":"Local model files only (HF_HUB_OFFLINE)."},"engine_ref":null,"profiles":[{"host":"ai-151","gpus":[1],"placement":"gpu1","min_warm_sec":15,"max_context":null}]},{"id":"qwen3-8b-q4","kind":"llamacpp","model":{"name":"Qwen3-8B (Qwen/Qwen3-8B-GGUF, Q4_K_M)","notes":"Small-LLM lane on the 1080 Ti via llama.cpp VULKAN backend (fleet's CUDA build excludes Pascal sm_61). 32K ctx + q8_0 KV ~8.1GiB."},"engine_ref":null,"profiles":[{"host":"ai-9","gpus":[0],"placement":"gpu0","min_warm_sec":30,"max_context":32768}]},{"id":"comfyui-flux-unchained","kind":"comfyui-workflow","model":{"name":"Flux Unchained (8-step hybrid) txt2img","notes":"API-format flux txt2img grounded in models present on ai-150 (fluxUnchained unet + t5xxl + ae vae). Derived from the saved flux_unchained_workflow.json. Replace `graph` with an exact ComfyUI 'Save (API Format)' export to match the UI workflow 1:1."},"engine_ref":"comfyui-engine","profiles":[]},{"id":"comfyui-map-sdxl","kind":"comfyui-workflow","model":{"family":"sdxl","purpose":"battlemap / sprite / fantasy-map image generation (the node that served the earlier sprite gens)"},"engine_ref":"comfyui-engine","profiles":[]},{"id":"comfyui-sprite-rgba","kind":"comfyui-workflow","model":{"family":"sdxl+layerdiffuse","purpose":"true RGBA transparent sprite generation (LayerDiffuse over SDXL)"},"engine_ref":"comfyui-engine","profiles":[]},{"id":"gemma-4-26b-a4b-q4","kind":"llamacpp","model":{"name":"gemma-4-26B-A4B-it","params":"26B (MoE, 4B active)","quant":"Q4_K_M","context":16384,"notes":"Wired from a previously-orphaned copy on ai-151. Offline tier (swap-in)."},"engine_ref":null,"profiles":[{"host":"ai-151","gpus":[0],"placement":"gpu0","min_warm_sec":90,"max_context":16384}]}],"by_capability":{"qwen3-30b-a3b-q4":[{"host":"ai-150","service":"qwen3-30b-a3b-q4__gpu0","fidelity":null,"placement":"gpu0","gpus":[0],"cpu_mib":0,"model":{"name":"Qwen3-30B-A3B-Instruct-2507","params":"30B (MoE, 3B active)","quant":"Q4_K_M","context":32768,"vision":false,"source":"https://huggingface.co/unsloth/Qwen3-30B-A3B-Instruct-2507-GGUF"}},{"host":"ai-151","service":"qwen3-30b-a3b-q4__gpu0","fidelity":null,"placement":"gpu0","gpus":[0],"cpu_mib":0,"model":{"name":"Qwen3-30B-A3B-Instruct-2507","params":"30B (MoE, 3B active)","quant":"Q4_K_M","context":32768,"vision":false,"source":"https://huggingface.co/unsloth/Qwen3-30B-A3B-Instruct-2507-GGUF"}}],"mistral-small-3.2-24b-q4":[{"host":"ai-150","service":"mistral-small-3.2-24b-q4__gpu1","fidelity":null,"placement":"gpu1","gpus":[1],"cpu_mib":0,"model":{"name":"Mistral-Small-3.2-24B-Instruct-2506","params":"24B","quant":"Q4_K_M","context":32768,"vision":false,"source":"https://huggingface.co/unsloth/Mistral-Small-3.2-24B-Instruct-2506-GGUF"}},{"host":"ai-151","service":"mistral-small-3.2-24b-q4__gpu0","fidelity":null,"placement":"gpu0","gpus":[0],"cpu_mib":0,"model":{"name":"Mistral-Small-3.2-24B-Instruct-2506","params":"24B","quant":"Q4_K_M","context":32768,"vision":false,"source":"https://huggingface.co/unsloth/Mistral-Small-3.2-24B-Instruct-2506-GGUF"}}],"qwen3-vl-32b-q8":[{"host":"ai-150","service":"qwen3-vl-32b-q8__all","fidelity":null,"placement":"all","gpus":[0,1],"cpu_mib":0,"model":{"name":"Qwen3-VL-32B-Instruct","params":"32B (dense)","quant":"Q8_0","context":49152,"vision":true,"source":"https://huggingface.co/unsloth/Qwen3-VL-32B-Instruct-GGUF","notes":"Offline deep-work tier. Spans both GPUs (tensor-split 60/40). Q8 weights + 48K ctx q8 KV. Higher capability than the live tier; evicted instantly when any live capability is requested."}}],"qwen3-vl-32b-q4":[{"host":"ai-150","service":"qwen3-vl-32b-q4__gpu0","fidelity":null,"placement":"gpu0","gpus":[0],"cpu_mib":0,"model":{"name":"Qwen3-VL-32B-Instruct","params":"32B","quant":"Q4_K_M","context":32768,"vision":true,"source":"https://huggingface.co/unsloth/Qwen3-VL-32B-Instruct-GGUF","notes":"Provisional vision-fast lane. Same model as deep_split but Q4 and gpu0-only, so it evicts the live director when used. Current use case (scene-image verify/caption) is bursty enough that this is fine; the eviction cost slots in between live calls. When portrait localization returns or extraction runs at scale, replace this with a genuinely small VLM (Qwen2-VL-2B / Moondream2 / SmolVLM2) on the .151 3090 so high-frequency cheap reads never touch the live hot path. Per the brief: vision-doc = low-frequency / high-stakes / fidelity-critical (32B is right); vision-fast = high-frequency / low-stakes / latency-sensitive (small + isolated lane is right)."}},{"host":"ai-151","service":"qwen3-vl-32b-q4__gpu0","fidelity":null,"placement":"gpu0","gpus":[0],"cpu_mib":0,"model":{"name":"Qwen3-VL-32B-Instruct","params":"32B","quant":"Q4_K_M","context":32768,"vision":true,"source":"https://huggingface.co/unsloth/Qwen3-VL-32B-Instruct-GGUF","notes":"Provisional vision-fast lane. Same model as deep_split but Q4 and gpu0-only, so it evicts the live director when used. Current use case (scene-image verify/caption) is bursty enough that this is fine; the eviction cost slots in between live calls. When portrait localization returns or extraction runs at scale, replace this with a genuinely small VLM (Qwen2-VL-2B / Moondream2 / SmolVLM2) on the .151 3090 so high-frequency cheap reads never touch the live hot path. Per the brief: vision-doc = low-frequency / high-stakes / fidelity-critical (32B is right); vision-fast = high-frequency / low-stakes / latency-sensitive (small + isolated lane is right)."}}],"comfyui-engine":[{"host":"ai-150","service":"comfyui-engine__gpu0","fidelity":null,"placement":"gpu0","gpus":[0],"cpu_mib":0,"model":{"name":"ComfyUI","notes":"Workflow-driven; loads checkpoints at runtime. flux1-dev-Q5_K_S + t5-xxl present in ComfyUI/models/. Burst tier: scene imagery queues until prose drains, never displaces live."}},{"host":"ai-150","service":"comfyui-engine__gpu1","fidelity":null,"placement":"gpu1","gpus":[1],"cpu_mib":0,"model":{"name":"ComfyUI","notes":"Workflow-driven; loads checkpoints at runtime. flux1-dev-Q5_K_S + t5-xxl present in ComfyUI/models/. Burst tier: scene imagery queues until prose drains, never displaces live."}},{"host":"ai-151","service":"comfyui-engine__gpu0","fidelity":null,"placement":"gpu0","gpus":[0],"cpu_mib":0,"model":{"name":"ComfyUI","notes":"Workflow-driven; loads checkpoints at runtime. flux1-dev-Q5_K_S + t5-xxl present in ComfyUI/models/. Burst tier: scene imagery queues until prose drains, never displaces live."}}],"comfyui-flux-unchained":[{"host":"ai-150","service":"comfyui-engine__gpu0","fidelity":null,"placement":"gpu0","gpus":[0],"cpu_mib":0,"model":{"name":"ComfyUI","notes":"Workflow-driven; loads checkpoints at runtime. flux1-dev-Q5_K_S + t5-xxl present in ComfyUI/models/. Burst tier: scene imagery queues until prose drains, never displaces live."}},{"host":"ai-150","service":"comfyui-engine__gpu1","fidelity":null,"placement":"gpu1","gpus":[1],"cpu_mib":0,"model":{"name":"ComfyUI","notes":"Workflow-driven; loads checkpoints at runtime. flux1-dev-Q5_K_S + t5-xxl present in ComfyUI/models/. Burst tier: scene imagery queues until prose drains, never displaces live."}},{"host":"ai-151","service":"comfyui-engine__gpu0","fidelity":null,"placement":"gpu0","gpus":[0],"cpu_mib":0,"model":{"name":"ComfyUI","notes":"Workflow-driven; loads checkpoints at runtime. flux1-dev-Q5_K_S + t5-xxl present in ComfyUI/models/. Burst tier: scene imagery queues until prose drains, never displaces live."}}],"comfyui-map-sdxl":[{"host":"ai-150","service":"comfyui-engine__gpu0","fidelity":null,"placement":"gpu0","gpus":[0],"cpu_mib":0,"model":{"name":"ComfyUI","notes":"Workflow-driven; loads checkpoints at runtime. flux1-dev-Q5_K_S + t5-xxl present in ComfyUI/models/. Burst tier: scene imagery queues until prose drains, never displaces live."}},{"host":"ai-150","service":"comfyui-engine__gpu1","fidelity":null,"placement":"gpu1","gpus":[1],"cpu_mib":0,"model":{"name":"ComfyUI","notes":"Workflow-driven; loads checkpoints at runtime. flux1-dev-Q5_K_S + t5-xxl present in ComfyUI/models/. Burst tier: scene imagery queues until prose drains, never displaces live."}},{"host":"ai-151","service":"comfyui-engine__gpu0","fidelity":null,"placement":"gpu0","gpus":[0],"cpu_mib":0,"model":{"name":"ComfyUI","notes":"Workflow-driven; loads checkpoints at runtime. flux1-dev-Q5_K_S + t5-xxl present in ComfyUI/models/. Burst tier: scene imagery queues until prose drains, never displaces live."}}],"comfyui-sprite-rgba":[{"host":"ai-150","service":"comfyui-engine__gpu0","fidelity":null,"placement":"gpu0","gpus":[0],"cpu_mib":0,"model":{"name":"ComfyUI","notes":"Workflow-driven; loads checkpoints at runtime. flux1-dev-Q5_K_S + t5-xxl present in ComfyUI/models/. Burst tier: scene imagery queues until prose drains, never displaces live."}},{"host":"ai-150","service":"comfyui-engine__gpu1","fidelity":null,"placement":"gpu1","gpus":[1],"cpu_mib":0,"model":{"name":"ComfyUI","notes":"Workflow-driven; loads checkpoints at runtime. flux1-dev-Q5_K_S + t5-xxl present in ComfyUI/models/. Burst tier: scene imagery queues until prose drains, never displaces live."}},{"host":"ai-151","service":"comfyui-engine__gpu0","fidelity":null,"placement":"gpu0","gpus":[0],"cpu_mib":0,"model":{"name":"ComfyUI","notes":"Workflow-driven; loads checkpoints at runtime. flux1-dev-Q5_K_S + t5-xxl present in ComfyUI/models/. Burst tier: scene imagery queues until prose drains, never displaces live."}}],"qwen3-embedding-4b-q8":[{"host":"ai-150","service":"qwen3-embedding-4b-q8__cpu","fidelity":null,"placement":"cpu","gpus":[],"cpu_mib":10000,"model":{"name":"Qwen3-Embedding-4B","params":"4B","quant":"Q8_0","context":32768,"vision":false,"source":"https://huggingface.co/Qwen/Qwen3-Embedding-4B-GGUF"}},{"host":"ai-151","service":"qwen3-embedding-4b-q8__gpu0","fidelity":null,"placement":"gpu0","gpus":[0],"cpu_mib":0,"model":{"name":"Qwen3-Embedding-4B","params":"4B","quant":"Q8_0","context":32768,"vision":false,"source":"https://huggingface.co/Qwen/Qwen3-Embedding-4B-GGUF"}}],"bge-reranker-v2-m3-q8":[{"host":"ai-150","service":"bge-reranker-v2-m3-q8__cpu","fidelity":null,"placement":"cpu","gpus":[],"cpu_mib":1500,"model":{"name":"BGE-reranker-v2-m3","params":"568M","quant":"Q8_0","context":8192,"vision":false,"source":"https://huggingface.co/gpustack/bge-reranker-v2-m3-GGUF"}}],"faster-whisper":[{"host":"ai-151","service":"faster-whisper__gpu1","fidelity":null,"placement":"gpu1","gpus":[1],"cpu_mib":0,"model":{"name":"faster-whisper @ Zoont/faster-whisper-large-v3-turbo-int8-ct2","notes":"Wrapped by stt_service (uvicorn app:app on the shared embedded python). Clients use `switchboard_client.transcribe(audio, ...)` — it handles ticket acquisition, multipart upload, and form fields. POST /transcribe accepts multipart `audio` + form fields (initial_prompt, language, model, word_timestamps); returns text + per-segment + per-word timing + no_speech/low_confidence flags. Default model swapped (Jun 5) from `large-v3-turbo` (runtime-quantized to int8_float16) to `Zoont/faster-whisper-large-v3-turbo-int8-ct2` (pre-quantized int8 weights from HuggingFace). Same Whisper model under the hood; pre-quantized weights load slightly faster and use less VRAM. Compute kept at int8_float16 for GPU mixed-precision inference. Default upload cap 10 MB (STT_MAX_UPLOAD_BYTES). Verified Jun 5: same transcription quality + sub-second inference on 2-6 s clips. STREAMING /transcribe_stream WebSocket added Jun 5 for real-time turn-taking (sb.transcribe_stream in client v1.2.0+). Server-side Silero VAD endpointing; partial+final events. Jun 5 fix: hotwords now passed via faster-whisper's native `hotwords=` (CTranslate2 logit biasing) instead of concatenated into initial_prompt -- the latter caused Whisper to echo the vocab list as the first partial of every utterance. Two backstops added: drop partials with no_speech_prob>threshold (cough/breath suppression), and _is_prompt_echo() check that suppresses any output text equal to the hotwords/prompt verbatim. Jun 5 followup: native hotwords still hallucinated the vocab list on unclear mic audio (DM AI). REPLACED the output-filter approach with TWO-PASS DECODE ARBITRATION in _decode(): when hotwords are configured, runs whisper twice (baseline=no-bias, biased=with hotwords) and arbitrates. Baseline empty/no_speech -> use baseline. Biased introduced hotwords baseline never saw -> use baseline (bias was artificial). Otherwise -> use biased (baseline confirms hotwords belong, biased just spells them better). Cost: 2x decode time per partial when hotwords configured. Source-level fix; legitimate one-word hotword answers (\"Karrthûn\" alone) pass through."}}],"kokoro":[{"host":"ai-151","service":"kokoro__gpu1","fidelity":null,"placement":"gpu1","gpus":[1],"cpu_mib":0,"model":{"name":"Kokoro-82M (hexgrad/Kokoro-82M, Apache-2.0)","notes":"Canonical TTS for the fleet (replaced AllTalk on cutover Jun 5). 82M-param StyleTTS2 distillation; ~190ms warm render for a 6.5s line (~33x realtime). 24kHz mono float32 output. Native `speed` parameter (0.5-2.0). 50+ voicepacks via /voices. Sidecar at C:\\AI\\kokoro_service installs the kokoro pip package into the existing C:\\AI\\python-3.12.10-embed-amd64 distribution (shared with agent + stt + diarization; torch 2.5.1+cu121 was already there). transformers pinned to <5.0 (5.x requires torch>=2.7). LANGUAGE EXTRAS: `kokoro` pulls `misaki[en]` only; non-English pipelines lazy-fail on import unless their misaki extras are installed too. install.ps1 explicitly pip installs `misaki[ja]`, `misaki[zh]`, `misaki[ko]`, and `ordered_set` to cover all currently-exposed non-English voicepacks (the Mandarin `lang=z` pipeline in particular imports `ordered_set` via `misaki.zh.transcription`). KPipeline init passes `repo_id='hexgrad/Kokoro-82M'` explicitly to suppress the noisy default-repo warning misaki logs per fresh language load. skip_if_missing keys on a .installed marker the install script touches at the end. Clients should use switchboard_client.tts_say* — direct /generate POSTs are supported but bypass the chunking and tag system."}}],"pyannote":[{"host":"ai-151","service":"pyannote__gpu1","fidelity":null,"placement":"gpu1","gpus":[1],"cpu_mib":0,"model":{"name":"pyannote-audio speaker-diarization-3.1","notes":"Local model files only (HF_HUB_OFFLINE)."}}],"gemma-4-26b-a4b-q4":[{"host":"ai-151","service":"gemma-4-26b-a4b-q4__gpu0","fidelity":null,"placement":"gpu0","gpus":[0],"cpu_mib":0,"model":{"name":"gemma-4-26B-A4B-it","params":"26B (MoE, 4B active)","quant":"Q4_K_M","context":16384,"notes":"Wired from a previously-orphaned copy on ai-151. Offline tier (swap-in)."}}],"qwen3-8b-q4":[{"host":"ai-9","service":"qwen3-8b-q4__gpu0","fidelity":null,"placement":"gpu0","gpus":[0],"cpu_mib":0,"model":{"name":"Qwen3-8B (Qwen/Qwen3-8B-GGUF, Q4_K_M)","notes":"Small-LLM lane on the 1080 Ti via llama.cpp VULKAN backend (fleet's CUDA build excludes Pascal sm_61). 32K ctx + q8_0 KV ~8.1GiB.","context":32768}}]}}