Instructions to use Youssofal/Qwen3.6-27B-MTPLX-Optimized with libraries, inference providers, notebooks, and local apps. Follow these links to get started.

Libraries

How to use Youssofal/Qwen3.6-27B-MTPLX-Optimized with MLX:

# Make sure mlx-lm is installed
# pip install --upgrade mlx-lm

# Generate text with mlx-lm
from mlx_lm import load, generate

model, tokenizer = load("Youssofal/Qwen3.6-27B-MTPLX-Optimized")

prompt = "Write a story about Einstein"
messages = [{"role": "user", "content": prompt}]
prompt = tokenizer.apply_chat_template(
    messages, add_generation_prompt=True
)

text = generate(model, tokenizer, prompt=prompt, verbose=True)

Notebooks
Google Colab
Kaggle
Local Apps
LM Studio

Pi new

How to use Youssofal/Qwen3.6-27B-MTPLX-Optimized with Pi:

Start the MLX server

# Install MLX LM:
uv tool install mlx-lm
# Start a local OpenAI-compatible server:
mlx_lm.server --model "Youssofal/Qwen3.6-27B-MTPLX-Optimized"

Configure the model in Pi

# Install Pi:
npm install -g @mariozechner/pi-coding-agent
# Add to ~/.pi/agent/models.json:
{
  "providers": {
    "mlx-lm": {
      "baseUrl": "http://localhost:8080/v1",
      "api": "openai-completions",
      "apiKey": "none",
      "models": [
        {
          "id": "Youssofal/Qwen3.6-27B-MTPLX-Optimized"
        }
      ]
    }
  }
}

Run Pi

# Start Pi in your project directory:
pi

Hermes Agent new

How to use Youssofal/Qwen3.6-27B-MTPLX-Optimized with Hermes Agent:

Start the MLX server

# Install MLX LM:
uv tool install mlx-lm
# Start a local OpenAI-compatible server:
mlx_lm.server --model "Youssofal/Qwen3.6-27B-MTPLX-Optimized"

Configure Hermes

# Install Hermes:
curl -fsSL https://hermes-agent.nousresearch.com/install.sh | bash
hermes setup
# Point Hermes at the local server:
hermes config set model.provider custom
hermes config set model.base_url http://127.0.0.1:8080/v1
hermes config set model.default Youssofal/Qwen3.6-27B-MTPLX-Optimized

Run Hermes

hermes

MLX LM

How to use Youssofal/Qwen3.6-27B-MTPLX-Optimized with MLX LM:

Generate or start a chat session

# Install MLX LM
uv tool install mlx-lm
# Interactive chat REPL
mlx_lm.chat --model "Youssofal/Qwen3.6-27B-MTPLX-Optimized"

Run an OpenAI-compatible server

# Install MLX LM
uv tool install mlx-lm
# Start the server
mlx_lm.server --model "Youssofal/Qwen3.6-27B-MTPLX-Optimized"
# Calling the OpenAI-compatible server with curl
curl -X POST "http://localhost:8000/v1/chat/completions" \
   -H "Content-Type: application/json" \
   --data '{
     "model": "Youssofal/Qwen3.6-27B-MTPLX-Optimized",
     "messages": [
       {"role": "user", "content": "Hello"}
     ]
   }'

Youssofal commited on 21 days ago

Commit

fe68f16

verified ·

1 Parent(s): bcdc6f3

Reframe model card: MTPLX coming soon, drop runtime tok/s claims

Browse files

Files changed (1) hide show

README.md +22 -67

README.md CHANGED Viewed

@@ -18,58 +18,37 @@ pipeline_tag: text-generation
 # Qwen3.6-27B MTPLX Optimized
-**Native MTP speculative decoding on Apple Silicon. 60+ tok/s on a 27B model at `temperature=0.6`.**
-This is the verified default checkpoint for [MTPLX](https://github.com/youssofal/mtplx) — an MLX-native inference engine that runs the model's own built-in Multi-Token-Prediction heads as a speculative drafter, with **mathematically exact** probability-ratio acceptance and residual correction.
-Not greedy-prefix-match. Not an external draft model. Not CUDA. The MTP draft path is exact at any temperature, so real coding settings (`temperature=0.6`, `top_p=0.95`, `top_k=20`) get the full speculative speedup *and* preserve the target model's output distribution.
-## Quick start
-```bash
-pip install mtplx        # or use the preview wheel on https://github.com/youssofal/mtplx
-mtplx quickstart         # interactive: pick model → mode → web/CLI, then chat
-```
-`mtplx quickstart` auto-downloads this model on first run if it isn't already on disk. It also detects existing copies in your local model folders (`~/models/`, LM Studio, HuggingFace cache) and lets you pick from them.
-To pin this model explicitly:
-```bash
-mtplx quickstart --model Youssofal/Qwen3.6-27B-MTPLX-Optimized
-```
-For the OpenAI-compatible API server only (no chat UI):
-```bash
-mtplx start --port 8000
-```
-The server then exposes:
-- `/v1/chat/completions` and `/v1/completions` (OpenAI-compatible)
-- `/v1/messages` (Anthropic-compatible, streaming SSE)
-- `/v1/models`, `/health`, `/metrics`
-Plug it into Open WebUI, Claude Code, Cline, Continue, or anything that speaks OpenAI.
 ## What's in this checkpoint
 | Component | Format |
 | --- | --- |
-| Trunk text + vision weights | MLX-affine mixed-precision, 8-bit Gated Delta Network linears, 4-bit MLP linears, BF16 norms |
 | MTP head sidecar (`mtp.safetensors`) | Calibrated CyanKiwi prequantized INT4 with BF16 MTP norms |
 | Vision encoder (`model-vision-*.safetensors`) | BF16, intact for multimodal use |
-| Runtime contract (`mtplx_runtime.json`) | Pins MTPLX version, arch, recommended profile, and verified hardware |
 | Tokenizer + chat template | Qwen3.6 vocabulary (248k tokens) |
-The MTP head is grafted from a separately calibrated INT4 sidecar (`Qwen3.6-27B-MTPLX-CyanKiwi-Packed-BF16-INT4-v3`) onto the MTPLX-specific GDN8-Speed4 trunk. This combination outperforms BF16 MTP on D2/D3/D4 acceptance under MTPLX's committed-history cache contract, contradicting the older assumption that BF16 MTP was strictly better than quantized MTP.
-## Acceptance numbers
-Per-position acceptance under MTPLX's `linear-gdn-from-conv-tape` verify path, depth 4, on `long_code` 192-token prompts at `temperature=0.6, top_p=0.95, top_k=20`:
-| Depth | MTPLX (this checkpoint) | vLLM MTP-5 oracle (3090, same temp) |
 | --- | --- | --- |
 | 1 | **97.62%** | 92.7% |
 | 2 | **95.24%** | 77.0% |
@@ -77,51 +56,27 @@ Per-position acceptance under MTPLX's `linear-gdn-from-conv-tape` verify path, d
 | 4 | **75.61%** | 50.9% |
 | 5 | — | 43.0% |
-MTPLX is higher at every depth than vLLM's MTP-5 implementation on the same Qwen3.6 family. Source: Phase 1 v4 oracle measurements, see [LOG.md](https://github.com/youssofal/mtplx/blob/main/LOG.md).
-## Performance
-Verified on Apple Silicon **M5 Max, 128 GB unified memory, macOS 26.3.1**, MLX 0.31.2:
-- **60.169 tok/s** clean-preflight headline on D3 / 192-token long-code prompts at `temp=0.6` — the cold cold-decode benchmark MTPLX was built around.
-- **2.54×** vs matched no-MTP autoregressive control (23.59 tok/s) on the same prompt and sampler.
-- **`max_diff = 0.0`** on the Phase 0H paged-verifier exactness gate (2048 ctx).
-- **D5 / 512-token long-code**: 43.871 tok/s with [97.09, 91.26, 82.35, 69.61, 58.82] per-position acceptance.
-Sustained no-fan long-context throughput is **lower** than the cold number — currently ~37 tok/s on 10k-token uncapped generation under macOS power-governor throttling. Closing that gap is the v0.2 deliverable; see the MTPLX repo's roadmap for the kernel-ladder plan. Until then, MTPLX ships three honest profiles:
-- `Safe` — ~37 tok/s steady, no fan changes, near-flat on long answers. Default for new users.
-- `Fast` — ~60 tok/s on short replies, decays on long ones because fans stay on Apple's default curve.
-- `Max` — Fast + ThermalForge fans pinned at 100%, sustained ~60 tok/s, loud. Opt-in via the wizard.
-## Hardware compatibility
-Tested on Apple Silicon M5 Max 128 GB. Should run on any Apple Silicon Mac with ≥ 24 GB unified memory, but only the M5 Max class is verified for the cold 60 tok/s number. M3/M4 numbers will be lower in proportion to memory bandwidth (the M5 Max has 614 GB/s).
-The model file footprint on disk is ~19 GB — most users will want at least 32 GB unified memory for comfortable batched inference.
 ## Provenance
 - **Base model**: [`Qwen/Qwen3.6-27B`](https://huggingface.co/Qwen/Qwen3.6-27B) (Apache 2.0).
 - **Quantization policy**: `mtplx-gdn8-speed4` — MLX-affine mixed-precision with uniform 8-bit GDN linears, 4-bit MLP, 4-bit `lm_head`, BF16 norms and the MTP head's `fc` projection.
 - **MTP sidecar**: `cyankiwi-calibrated-int4-prequantized`, calibrated separately with MLX-affine quantization and grafted onto the GDN8-Speed4 trunk.
-- **Runtime contract**: `mtplx_runtime.json` pins the architecture (`qwen3-next-mtp`), recommended profile (`stable`), and exactness baseline (`max_diff = 0.0` on Phase 0H paged-verifier smoke at 2048 ctx).
 ## Limitations
-- **Apple Silicon only** at present. MTPLX uses MLX as its primary backend; CUDA / x86 are not supported in v0.1.
-- **Cold and steady-state are different numbers**. The 60+ tok/s figure is the cold long-code D3 path. Sustained no-fan long-context throughput is lower until v0.2 lands the kernel ladder.
-- **Verified runtime is Qwen3.6-27B**. MTPLX recognizes other MTP architectures (DeepSeek V3 MTP, GLM4 MoE MTP, MiMo, MiniMax M2 MTP, etc.) and will report them via `mtplx inspect <repo>`, but only Qwen3-Next-class artifacts have the verified runtime contract. Other architectures are detected as `architecture-compatible-but-unverified` and require `--unsafe-force-unverified --yes` to run.
-- **Not greedy-only**. The whole point is exactness at `temperature=0.6`. If you only ever decode greedily, simpler tools like `mlx-lm` or the upstream Qwen MLX path are smaller dependencies.
 ## License
-This checkpoint is released under the **Apache License 2.0**, matching the Qwen3.6-27B base model. The MTPLX runtime is also Apache 2.0.
 ## Citation
-If MTPLX or this checkpoint helped your work, please cite:
 ```bibtex
 @misc{mtplx2026,
   author       = {Youssof Al},
@@ -133,5 +88,5 @@ If MTPLX or this checkpoint helped your work, please cite:
 ## Links
-- **Runtime**: [github.com/youssofal/mtplx](https://github.com/youssofal/mtplx)
 - **Base model**: [Qwen/Qwen3.6-27B](https://huggingface.co/Qwen/Qwen3.6-27B)

 # Qwen3.6-27B MTPLX Optimized
+> ## MTPLX — coming soon
+>
+> This checkpoint is the verified default for the upcoming **MTPLX** inference engine — an MLX-native runtime for **native Multi-Token-Prediction speculative decoding on Apple Silicon**. **The runtime is not publicly released yet.** This model card is published in advance so you can review the architecture, MTP head, and quantization decisions while the runtime is finalized. Watch [github.com/youssofal/mtplx](https://github.com/youssofal/mtplx) for the release.
+This artifact pairs the Qwen3.6-27B trunk — MLX-quantized with MTPLX's `gdn8-speed4` policy (8-bit Gated Delta Network linears, 4-bit MLP, BF16 norms) — with a **calibrated INT4 Multi-Token-Prediction sidecar** grafted onto the trunk. The MTP head is what enables *native* speculative decoding: the model drafts its own tokens, with no external draft model required.
+When MTPLX is released, it will accept those draft tokens with **mathematically exact** probability-ratio acceptance and residual correction, so the speculative path stays distribution-preserving at realistic coding settings (`temperature=0.6`, `top_p=0.95`, `top_k=20`) — not just greedy.
+Until then you can still:
+- Inspect the architecture and MTP tensors with any `safetensors` reader.
+- Use the trunk weights with [`mlx-lm`](https://github.com/ml-explore/mlx-lm) for ordinary autoregressive decoding (the MTP head is sidecar-only and ignored by `mlx-lm`).
+- Read the calibration / quantization metadata in `mtplx_runtime.json` and `config.json` to understand the build.
 ## What's in this checkpoint
 | Component | Format |
 | --- | --- |
+| Trunk text + vision weights | MLX-affine mixed-precision: 8-bit Gated Delta Network linears, 4-bit MLP linears, BF16 norms |
 | MTP head sidecar (`mtp.safetensors`) | Calibrated CyanKiwi prequantized INT4 with BF16 MTP norms |
 | Vision encoder (`model-vision-*.safetensors`) | BF16, intact for multimodal use |
+| Runtime contract (`mtplx_runtime.json`) | Pins architecture, recommended profile, and exactness baseline |
 | Tokenizer + chat template | Qwen3.6 vocabulary (248k tokens) |
+The MTP head is grafted from a separately calibrated INT4 sidecar (`Qwen3.6-27B-MTPLX-CyanKiwi-Packed-BF16-INT4-v3`) onto the MTPLX-specific GDN8-Speed4 trunk. This combination outperforms BF16 MTP on D2/D3/D4 acceptance under MTPLX's committed-history cache contract.
+## MTP draft acceptance
+These numbers describe the **MTP head's draft quality** — a property of the model itself, independent of any runtime's wall-clock throughput. Per-position acceptance under exact probability-ratio sampling at `temperature=0.6, top_p=0.95, top_k=20`:
+| Depth | This checkpoint | vLLM MTP-5 oracle (3090, same temp) |
 | --- | --- | --- |
 | 1 | **97.62%** | 92.7% |
 | 2 | **95.24%** | 77.0% |
 | 4 | **75.61%** | 50.9% |
 | 5 | — | 43.0% |
+Higher acceptance at every depth than vLLM's MTP-5 implementation on the same Qwen3.6 family, measured on `long_code` 192-token prompts.
 ## Provenance
 - **Base model**: [`Qwen/Qwen3.6-27B`](https://huggingface.co/Qwen/Qwen3.6-27B) (Apache 2.0).
 - **Quantization policy**: `mtplx-gdn8-speed4` — MLX-affine mixed-precision with uniform 8-bit GDN linears, 4-bit MLP, 4-bit `lm_head`, BF16 norms and the MTP head's `fc` projection.
 - **MTP sidecar**: `cyankiwi-calibrated-int4-prequantized`, calibrated separately with MLX-affine quantization and grafted onto the GDN8-Speed4 trunk.
+- **Runtime contract**: `mtplx_runtime.json` pins the architecture (`qwen3-next-mtp`), recommended profile, and exactness baseline.
 ## Limitations
+- **The MTPLX runtime is not yet released.** Without it, you can still use the trunk weights with `mlx-lm` for ordinary AR decoding — but the MTP draft path that this checkpoint was built for requires MTPLX.
+- **Apple Silicon focus.** MTPLX targets MLX as its primary backend; CUDA / x86 are not supported.
+- **Verified architecture is Qwen3-Next.** MTPLX recognizes other MTP architectures (DeepSeek V3 MTP, GLM4 MoE MTP, MiMo, MiniMax M2 MTP, etc.) but only Qwen3-Next-class artifacts have a verified runtime contract today.
 ## License
+This checkpoint is released under the **Apache License 2.0**, matching the Qwen3.6-27B base model.
 ## Citation
 ```bibtex
 @misc{mtplx2026,
   author       = {Youssof Al},
 ## Links
+- **Runtime (coming soon)**: [github.com/youssofal/mtplx](https://github.com/youssofal/mtplx)
 - **Base model**: [Qwen/Qwen3.6-27B](https://huggingface.co/Qwen/Qwen3.6-27B)