Source-ready MLX Gradio demo

Qwen3.8-27B
Uncensored MLX 4-bit

A multimodal local demo for the gated orcarouter model. The interactive Gradio application and local installation helpers are included in this repository.

Hosted inference is not active yet. The 4-bit MLX model needs a dedicated CUDA GPU for reliable Hugging Face hosting. This free Static Space keeps the implementation public and ready while GPU access is pending. The local Apple Silicon path remains available now.

Local 4-bit model

Install the 4-bit weights under ~/.ollama/models/Qwen3.8-27B-Uncensored-MLX/4-bit.

Run the UI

Execute ./setup_local.sh, then ./run_local.sh. The local UI opens at port 7860.

Hosted upgrade path

Switch the Space front matter to Gradio, add HF_TOKEN, and attach a persistent GPU volume.

Repository contents

  1. app.py - streaming text and image chat UI.
  2. download_4bit.sh - gated 4-bit snapshot download helper.
  3. setup_local.sh and run_local.sh - Apple Silicon setup and launch scripts.