Local MLX Vision vs Text Brain: Separate Ports or Fake Eyes
DAJAI Stewart
Why vision and text must not share a pretend endpoint—Apple Silicon MLX vision for captioning and compliance gates, structured JSON outputs, KeepAlive, memory pressure.
If your "vision model" is a text-only LLM on the same port, you do not have eyes. You have a chatbot with confidence.
Split the specialists
| Role | Surface | Job |
|---|---|---|
| Text brain | local reasoning port | decide, plan, score |
| Vision | MLX vision port | caption, tag, age-gate style JSON |
| Embeddings | small sidecar | search vectors only |