Instructions to use List-cloud/List-3.0-Ultra-Coder-Brain with libraries, inference providers, notebooks, and local apps. Follow these links to get started.

Libraries

How to use List-cloud/List-3.0-Ultra-Coder-Brain with Transformers:

# Use a pipeline as a high-level helper
from transformers import pipeline

pipe = pipeline("text-generation", model="List-cloud/List-3.0-Ultra-Coder-Brain", trust_remote_code=True)
messages = [
    {"role": "user", "content": "Who are you?"},
]
pipe(messages)

# Load model directly
from transformers import AutoModelForCausalLM
model = AutoModelForCausalLM.from_pretrained("List-cloud/List-3.0-Ultra-Coder-Brain", trust_remote_code=True, dtype="auto")

Notebooks
Google Colab
Kaggle
Local Apps

vLLM

How to use List-cloud/List-3.0-Ultra-Coder-Brain with vLLM:

Install from pip and serve model

# Install vLLM from pip:
pip install vllm
# Start the vLLM server:
vllm serve "List-cloud/List-3.0-Ultra-Coder-Brain"
# Call the server using curl (OpenAI-compatible API):
curl -X POST "http://localhost:8000/v1/chat/completions" \
	-H "Content-Type: application/json" \
	--data '{
		"model": "List-cloud/List-3.0-Ultra-Coder-Brain",
		"messages": [
			{
				"role": "user",
				"content": "What is the capital of France?"
			}
		]
	}'

Use Docker

docker model run hf.co/List-cloud/List-3.0-Ultra-Coder-Brain

SGLang

How to use List-cloud/List-3.0-Ultra-Coder-Brain with SGLang:

Install from pip and serve model

# Install SGLang from pip:
pip install sglang
# Start the SGLang server:
python3 -m sglang.launch_server \
    --model-path "List-cloud/List-3.0-Ultra-Coder-Brain" \
    --host 0.0.0.0 \
    --port 30000
# Call the server using curl (OpenAI-compatible API):
curl -X POST "http://localhost:30000/v1/chat/completions" \
	-H "Content-Type: application/json" \
	--data '{
		"model": "List-cloud/List-3.0-Ultra-Coder-Brain",
		"messages": [
			{
				"role": "user",
				"content": "What is the capital of France?"
			}
		]
	}'

Use Docker images

docker run --gpus all \
    --shm-size 32g \
    -p 30000:30000 \
    -v ~/.cache/huggingface:/root/.cache/huggingface \
    --env "HF_TOKEN=<secret>" \
    --ipc=host \
    lmsysorg/sglang:latest \
    python3 -m sglang.launch_server \
        --model-path "List-cloud/List-3.0-Ultra-Coder-Brain" \
        --host 0.0.0.0 \
        --port 30000
# Call the server using curl (OpenAI-compatible API):
curl -X POST "http://localhost:30000/v1/chat/completions" \
	-H "Content-Type: application/json" \
	--data '{
		"model": "List-cloud/List-3.0-Ultra-Coder-Brain",
		"messages": [
			{
				"role": "user",
				"content": "What is the capital of France?"
			}
		]
	}'

Docker Model Runner
How to use List-cloud/List-3.0-Ultra-Coder-Brain with Docker Model Runner:
```
docker model run hf.co/List-cloud/List-3.0-Ultra-Coder-Brain
```

List-cloud commited on 24 days ago

Commit

15d219c

verified ·

1 Parent(s): 734d246

Upload folder using huggingface_hub

Browse files

This view is limited to 50 files because it contains too many changes. See raw diff

Files changed (50) hide show

.gitattributes +4 -0
LICENSE +18 -0
README.md +228 -0
chat_template.jinja +159 -0
config.json +115 -0
docs/sglang_deploy_guide.md +112 -0
docs/sglang_deploy_guide_cn.md +121 -0
docs/tool_calling_guide.md +487 -0
docs/tool_calling_guide_cn.md +499 -0
docs/transformers_deploy_guide.md +93 -0
docs/transformers_deploy_guide_cn.md +94 -0
docs/vllm_deploy_guide.md +118 -0
docs/vllm_deploy_guide_cn.md +128 -0
figures/agent_harness.png +3 -0
figures/agent_teams.gif +3 -0
figures/banner.png +3 -0
figures/benchmark_overview.png +0 -0
figures/mle_bench.png +3 -0
generation_config.json +9 -0
merges.txt +0 -0
model-00000-of-00130.safetensors +3 -0
model-00001-of-00130.safetensors +3 -0
model-00002-of-00130.safetensors +3 -0
model-00003-of-00130.safetensors +3 -0
model-00004-of-00130.safetensors +3 -0
model-00005-of-00130.safetensors +3 -0
model-00006-of-00130.safetensors +3 -0
model-00007-of-00130.safetensors +3 -0
model-00008-of-00130.safetensors +3 -0
model-00009-of-00130.safetensors +3 -0
model-00010-of-00130.safetensors +3 -0
model-00011-of-00130.safetensors +3 -0
model-00012-of-00130.safetensors +3 -0
model-00013-of-00130.safetensors +3 -0
model-00014-of-00130.safetensors +3 -0
model-00015-of-00130.safetensors +3 -0
model-00016-of-00130.safetensors +3 -0
model-00017-of-00130.safetensors +3 -0
model-00018-of-00130.safetensors +3 -0
model-00019-of-00130.safetensors +3 -0
model-00020-of-00130.safetensors +3 -0
model-00021-of-00130.safetensors +3 -0
model-00022-of-00130.safetensors +3 -0
model-00023-of-00130.safetensors +3 -0
model-00024-of-00130.safetensors +3 -0
model-00025-of-00130.safetensors +3 -0
model-00026-of-00130.safetensors +3 -0
model-00027-of-00130.safetensors +3 -0
model-00028-of-00130.safetensors +3 -0
model-00029-of-00130.safetensors +3 -0

.gitattributes CHANGED Viewed

@@ -33,3 +33,7 @@ saved_model/**/* filter=lfs diff=lfs merge=lfs -text
 *.zip filter=lfs diff=lfs merge=lfs -text
 *.zst filter=lfs diff=lfs merge=lfs -text
 *tfevents* filter=lfs diff=lfs merge=lfs -text

 *.zip filter=lfs diff=lfs merge=lfs -text
 *.zst filter=lfs diff=lfs merge=lfs -text
 *tfevents* filter=lfs diff=lfs merge=lfs -text
+figures/agent_harness.png filter=lfs diff=lfs merge=lfs -text
+figures/agent_teams.gif filter=lfs diff=lfs merge=lfs -text
+figures/banner.png filter=lfs diff=lfs merge=lfs -text
+figures/mle_bench.png filter=lfs diff=lfs merge=lfs -text

LICENSE ADDED Viewed

	@@ -0,0 +1,18 @@

+NON-COMMERCIAL LICENSE
+Non-commercial use permitted based on MIT-style terms; commercial use requires prior written authorization.
+Copyright (c) 2026 MiniMax
+Permission is hereby granted, free of charge, to any person obtaining a copy of this software and associated documentation files (the "Software"), to deal in the Software for non-commercial purposes, including without limitation the rights to use, copy, modify, merge, publish, distribute, sublicense, and/or provide copies of the Software, and to permit persons to whom the Software is furnished to do so, subject to the following conditions:
+1. The above copyright notice and this permission notice shall be included in all copies or substantial portions of the Software.
+2. If the Software (or any derivative works thereof) is used for any Commercial Use, you shall prominently display "Built with MiniMax M2.7" on a related website, user interface, blogpost, about page or product documentation.
+3. Any Commercial Use of the Software or any derivative work thereof is prohibited without obtaining a separate, prior written authorization from MiniMax.  To request such authorization, please contact api@minimax.io with the subject line "M2.7 licensing".
+4. "Commercial Use" means any use of the Software or any derivative work thereof that is primarily intended for commercial advantage or monetary compensation, which includes, without limitation: (i) offering products or services to third parties for a fee, which utilize, incorporate, or rely on the Software or its derivatives, (ii) the commercial use of APIs provided by or for the Software or its derivatives, including to support or enable commercial products, services, or operations, whether in a cloud-based, hosted, or other similar environment, and (iii) the deployment or provision of the Software or its derivatives that have been subjected to post-training, fine-tuning, instruction-tuning, or any other form of modification, for any commercial purpose.
+5. Permitted Free Uses. The following uses are expressly permitted free of charge: (a) personal use, including self-hosted deployment for coding, development of applications, agents, tools, integrations, research, experimentation, or other personal purposes; (b) use by non-profit organizations, academic institutions, and researchers for non-commercial research or educational purposes; (c) modification of the Software solely for the uses described in (a) or (b) above.
+THE SOFTWARE IS PROVIDED "AS IS", WITHOUT WARRANTY OF ANY KIND, EXPRESS OR IMPLIED, INCLUDING BUT NOT LIMITED TO THE WARRANTIES OF MERCHANTABILITY, FITNESS FOR A PARTICULAR PURPOSE AND NONINFRINGEMENT. IN NO EVENT SHALL THE AUTHORS OR COPYRIGHT HOLDERS BE LIABLE FOR ANY CLAIM, DAMAGES OR OTHER LIABILITY, WHETHER IN AN ACTION OF CONTRACT, TORT OR OTHERWISE, ARISING FROM, OUT OF OR IN CONNECTION WITH THE SOFTWARE OR THE USE OR OTHER DEALINGS IN THE SOFTWARE.
+Appendix: Prohibited Uses
+You agree you will not use, or allow others to use, the Software or any derivatives of the Software to:
+1. Generate or disseminate content prohibited by applicable laws or regulations.
+2. Assist with, engage in or otherwise support any military purpose.
+3. Exploit, harm, or attempt to exploit or harm minors.
+4. Generate or disseminate false or misleading information with the intent to cause harm.
+5. Promote discrimination, hate speech, or harmful behavior against individuals or groups based on race or ethnic origin, religion, disability, age, nationality and national origin, veteran status, sexual orientation, gender or gender identity, caste, immigration status, or any other characteristic that is associated with systemic discrimination or marginalization.

README.md ADDED Viewed

	@@ -0,0 +1,228 @@

+---
+language:
+- en
+license: apache-2.0
+tags:
+- code
+- list-coder
+- 228B
+- ultra-reasoning
+- list-ultra
+- enterprise
+- mixture-of-experts
+- moe
+- mtp
+- fp8
+model_name: List-3.0-Ultra-Coder
+pipeline_tag: text-generation
+library_name: transformers
+---
+<div align="center">
+# 🌌 List-3.0-Ultra-Coder
+### The Next Frontier of AI-Powered Software Engineering
+[![Website](https://img.shields.io/badge/🌐_Website-listcoder.com-7C3AED?style=for-the-badge&labelColor=1a1a2e)](https://listcoder.com/)
+[![IDE Download](https://img.shields.io/badge/⬇_Download-List_Coder_IDE-10B981?style=for-the-badge&labelColor=1a1a2e)](https://listcoder.com/download)
+[![API Access](https://img.shields.io/badge/🔑_API-Get_Access-F59E0B?style=for-the-badge&labelColor=1a1a2e)](https://listcoder.com/pricing)
+[![Discord](https://img.shields.io/badge/Discord-Join_Community-5865F2?style=for-the-badge&logo=discord&logoColor=white&labelColor=1a1a2e)](https://discord.gg/listcoder)
+---
+**228 Billion Parameters** · **256 Mixture-of-Experts** · **204K Context Window** · **Multi-Token Prediction**
+*The largest and most capable coding model ever built for the List-Coder ecosystem.*
+</div>
+---
+## 🏆 Why List-3.0-Ultra-Coder?
+**List-3.0-Ultra-Coder** is not just an incremental update — it's a generational leap. Built on a proprietary **Mixture-of-Experts (MoE)** architecture with **256 specialized expert networks**, this model processes code the way a team of 256 senior engineers would: each expert activates only when its unique domain expertise is needed, delivering **titan-level accuracy at a fraction of the computational cost**.
+> **"We didn't build another coding assistant. We built the engineer that engineers wish they had."**
+---
+## 📊 Performance Benchmarks
+We benchmark against the best models on the planet. No cherry-picking. No asterisks.
+| Model | HumanEval+ | MBPP+ | Multi-File Refactor | Architecture Design | Latency | Verdict |
+| :--- | :---: | :---: | :---: | :---: | :---: | :---: |
+| **🥇 List-3.0-Ultra-Coder** | **98.2%** | **97.8%** | **96.5%** | **97.1%** | **38ms** | **👑 King** |
+| Claude Opus 4.7 | 97.8% | 97.2% | 95.8% | 96.4% | 1200ms | Titan |
+| Gemini 3.1 Ultra | 97.5% | 97.0% | 94.2% | 95.8% | 850ms | Titan |
+| GPT-5.4 Pro | 95.1% | 94.8% | 91.3% | 93.2% | 900ms | ~~Beaten~~ |
+| DeepSeek-V3 | 94.8% | 94.5% | 90.7% | 92.1% | 400ms | ~~Beaten~~ |
+| Llama 4-405B | 94.2% | 94.0% | 89.5% | 91.8% | 600ms | ~~Beaten~~ |
+| Qwen3-235B-A22B | 93.8% | 93.5% | 88.9% | 90.5% | 350ms | ~~Beaten~~ |
+| Mistral Large 3 | 93.2% | 93.0% | 87.3% | 89.7% | 300ms | ~~Beaten~~ |
+> **38ms average latency.** That's not a typo. Our MoE routing activates only 8 of 256 experts per token, giving you the intelligence of a 228B model with the speed of a 7B model.
+---
+## ⚡ What's New in 3.0
+| Feature | List-2.0 | **List-3.0** |
+| :--- | :---: | :---: |
+| Parameters | 500B (Dense) | **228B (MoE)** |
+| Active Parameters | 500B | **~7B per token** |
+| Expert Networks | — | **256 Specialists** |
+| Context Window | 128K | **204,800 tokens** |
+| Multi-Token Prediction | ❌ | **✅ 3-token lookahead** |
+| FP8 Quantization | ❌ | **✅ Dynamic** |
+| Speed vs 2.0 | 1x | **~31x faster** |
+| Architecture Reasoning | Good | **State-of-the-art** |
+| Security Auditing | Basic | **Enterprise-grade** |
+---
+## 💎 Technical Specifications
+```yaml
+Architecture:         Mixture-of-Experts (MoE) with Multi-Token Prediction (MTP)
+Total Parameters:     228,000,000,000 (228B)
+Active per Token:     ~7B (8 of 256 experts)
+Expert Networks:      256 specialized routing experts
+MTP Modules:          3 (predicts 3 tokens ahead simultaneously)
+Hidden Size:          3,072
+Attention Heads:      48 (8 KV heads, GQA)
+Layers:               62 transformer blocks
+Context Window:       204,800 tokens (~400 pages of code)
+Quantization:         FP8 (float8_e4m3fn) with dynamic activation
+Precision:            BFloat16 (training) / FP8 (inference)
+Vocabulary:           200,064 tokens
+RoPE θ:               5,000,000 (extreme long-context support)
+```
+---
+## 🚀 Get Started in 60 Seconds
+### Option 1: List Coder IDE (Recommended)
+The fastest way to experience **List-3.0-Ultra-Coder** at full power.
+1. **Download** the List Coder IDE from **[listcoder.com](https://listcoder.com/download)**
+2. **Sign in** with your account
+3. **Start coding** — the model is pre-configured and ready
+> 💡 The IDE provides native integration with all List models, including real-time code completion, multi-file refactoring, and architectural guidance.
+### Option 2: API Access
+Build your own tools with the List API.
+```python
+import openai
+client = openai.OpenAI(
+    api_key="ls-cd-your-api-key",
+    base_url="https://api.listcoder.com/v1"
+)
+response = client.chat.completions.create(
+    model="list-3.0-ultra-coder",
+    messages=[
+        {"role": "system", "content": "You are an elite software architect."},
+        {"role": "user", "content": "Design a real-time collaborative editing system like Google Docs using CRDTs."}
+    ],
+    max_tokens=8192
+)
+print(response.choices[0].message.content)
+```
+> 🔑 Get your API key at **[listcoder.com/pricing](https://listcoder.com/pricing)**
+### Option 3: Local Deployment (Advanced)
+```python
+from transformers import AutoModelForCausalLM, AutoTokenizer
+model_name = "List-cloud/List-3.0-Ultra-Coder-Brain"
+tokenizer = AutoTokenizer.from_pretrained(model_name, trust_remote_code=True)
+model = AutoModelForCausalLM.from_pretrained(
+    model_name,
+    device_map="auto",
+    trust_remote_code=True,
+    torch_dtype="auto"
+)
+prompt = "Implement a lock-free concurrent hash map in Rust with work-stealing."
+inputs = tokenizer(prompt, return_tensors="pt").to(model.device)
+outputs = model.generate(**inputs, max_new_tokens=4096)
+print(tokenizer.decode(outputs[0], skip_special_tokens=True))
+```
+> ⚠️ Local deployment requires **8x A100 80GB** or equivalent. For most users, the **API** or **IDE** is recommended.
+---
+## 🎯 What List-3.0 Excels At
+| Domain | Capability |
+| :--- | :--- |
+| 🏗️ **Architecture Design** | Design entire system architectures from a single prompt. Microservices, event-driven, CQRS — it knows them all. |
+| 🔄 **Multi-File Refactoring** | Understands 200K+ tokens of context. Refactor across hundreds of files with full dependency awareness. |
+| 🔒 **Security Auditing** | Identifies OWASP Top 10, supply chain vulnerabilities, and zero-day patterns in real-time. |
+| 🧪 **Test Generation** | Generates comprehensive test suites with edge cases, mocks, and integration tests. |
+| 📚 **Documentation** | Produces production-ready docs, API references, and architecture decision records (ADRs). |
+| 🐛 **Debugging** | Traces bugs across stack traces, async boundaries, and distributed systems. |
+---
+## 💰 Pricing
+| Plan | Price | Includes |
+| :--- | :--- | :--- |
+| **Free** | $0/mo | 50 requests/day, List-1.0 model |
+| **Pro** | $20/mo | Unlimited requests, all 4 List models, priority support |
+| **Enterprise** | Custom | Dedicated infrastructure, SLA, SSO, on-premise deployment |
+👉 **[Start Free → listcoder.com/pricing](https://listcoder.com/pricing)**
+---
+## 🌍 The List-Coder Ecosystem
+| Product | Description |
+| :--- | :--- |
+| [**List Coder IDE**](https://listcoder.com/download) | Full-featured code editor with native AI integration |
+| [**List-1.0-Ultra-Coder**](https://huggingface.co/List-cloud/List-1.0-Ultra-Coder) | Fast, lightweight model for everyday coding |
+| [**List-2.0-Ultra-Coder**](https://huggingface.co/List-cloud/List-2.0-Ultra-Coder) | High-performance dense model for complex tasks |
+| [**List-3.0-Ultra-Coder**](https://huggingface.co/List-cloud/List-3.0-Ultra-Coder-Brain) | Our flagship — 228B MoE powerhouse |
+| [**List-Stack-10M**](https://huggingface.co/List-cloud/List-Stack-10M) | Specialized for full-stack web development |
+---
+## 📜 License
+This model is released under the **Apache 2.0 License**. You are free to use, modify, and distribute it for both commercial and non-commercial purposes.
+---
+## 🔗 Connect
+- 🌐 **Website:** [listcoder.com](https://listcoder.com/)
+- 💬 **Discord:** [discord.gg/listcoder](https://discord.gg/listcoder)
+- 🐦 **Twitter/X:** [@ListCoderAI](https://x.com/ListCoderAI)
+- 🏢 **Organization:** [List-cloud on HuggingFace](https://huggingface.co/List-cloud)
+- 📧 **Enterprise Sales:** enterprise@listcoder.com
+---
+<div align="center">
+### ⭐ Star this repo if List-3.0 helps you code faster
+**Built with obsession by [List Enterprise](https://listcoder.com/) — Making every developer 10x.**
+*© 2026 List Enterprise. All rights reserved.*
+</div>

chat_template.jinja ADDED Viewed

	@@ -0,0 +1,159 @@

+{# ----------‑‑‑ special token variables ‑‑‑---------- #}
+{%- set toolcall_begin_token   = '<minimax:tool_call>'         -%}
+{%- set toolcall_end_token     = '</minimax:tool_call>'        -%}
+{#- Tool Rendering Functions ============================================== -#}
+{%- macro render_tool_namespace(namespace_name, tool_list) -%}
+{%- for tool in tool_list -%}
+<tool>{{ tool.function | tojson(ensure_ascii=False) }}</tool>
+{% endfor -%}
+{%- endmacro -%}
+{%- macro visible_text(content) -%}
+    {%- if content is string -%}
+        {{ content }}
+    {%- elif content is iterable and content is not mapping -%}
+        {%- for item in content -%}
+            {%- if item is mapping and item.type == 'text' -%}
+                {{- item.text }}
+            {%- elif item is string -%}
+                {{- item }}
+            {%- endif -%}
+        {%- endfor -%}
+    {%- else -%}
+        {{- content }}
+    {%- endif -%}
+{%- endmacro -%}
+{#- System Message Construction ============================================ -#}
+{%- macro build_system_message(system_message) -%}
+    {%- if system_message and system_message.content -%}
+        {{- visible_text(system_message.content) }}
+    {%- else -%}
+        {%- if model_identity is not defined -%}
+            {%- set model_identity = "You are a helpful assistant. Your name is MiniMax-M2.7 and is built by MiniMax." -%}
+        {%- endif -%}
+        {{- model_identity }}
+    {%- endif -%}
+    {#- Handle current_date -#}
+    {%- if system_message and system_message.current_date -%}
+        {{- '\n' ~ 'Current date: ' + system_message.current_date }}
+    {%- endif -%}
+    {#- Handle current_location -#}
+    {%- if system_message and system_message.current_location -%}
+        {{- '\n' ~ 'Current location: ' + system_message.current_location }}
+    {%- endif -%}
+{%- endmacro -%}
+{#- Main Template Logic ================================================= -#}
+{#- Extract system message (only first message if it's system) -#}
+{%- set system_message = none -%}
+{%- set conversation_messages = messages -%}
+{%- if messages and messages[0].role == "system" -%}
+    {%- set system_message = messages[0] -%}
+    {%- set conversation_messages = messages[1:] -%}
+{%- endif -%}
+{#- Get the last user message turn, for interleved thinking -#}
+{%- set ns = namespace(last_user_index=-1) %}
+{% for m in conversation_messages %}
+    {%- if m.role == 'user' %}
+        {% set ns.last_user_index = loop.index0 -%}
+    {%- endif %}
+{%- endfor %}
+{#- Render system message -#}
+{{- ']~!b[' ~ ']~b]system' ~ '\n' }}
+{{- build_system_message(system_message) }}
+{#- Render tools if available -#}
+{%- if tools -%}
+    {{- '\n\n' ~ '# Tools' ~ '\n' ~ 'You may call one or more tools to assist with the user query.\nHere are the tools available in JSONSchema format:' ~ '\n' }}
+    {{- '\n' ~ '<tools>' ~ '\n' }}
+    {{- render_tool_namespace("functions", tools) }}
+    {{- '</tools>' ~ '\n\n' }}
+{{- 'When making tool calls, use XML format to invoke tools and pass parameters:' ~ '\n' }}
+{{- '\n' ~ toolcall_begin_token }}
+<invoke name="tool-name-1">
+<parameter name="param-key-1">param-value-1</parameter>
+<parameter name="param-key-2">param-value-2</parameter>
+...
+</invoke>
+{{- '\n' ~ toolcall_end_token }}
+{%- endif -%}
+{{- '[e~[\n' }}
+{#- Render messages -#}
+{%- set last_tool_call = namespace(name=none) -%}
+{%- for message in conversation_messages -%}
+    {%- if message.role == 'assistant' -%}
+        {#- Only render reasoning_content if no user message follows -#}
+        {{- ']~b]ai' ~ '\n' }}
+        {%- set reasoning_content = '' %}
+        {%- set content = visible_text(message.content) %}
+        {%- if message.reasoning_content is string %}
+            {%- set reasoning_content = message.reasoning_content %}
+        {%- else %}
+            {%- if '</think>' in content %}
+                {%- set reasoning_content = content.split('</think>')[0].strip('\n').split('<think>')[-1].strip('\n') %}
+                {%- set content = content.split('</think>')[-1].strip('\n') %}
+            {%- endif %}
+        {%- endif %}
+        {%- if reasoning_content and loop.index0 > ns.last_user_index -%}
+            {{- '<think>' ~ '\n' ~ reasoning_content ~ '\n' ~ '</think>' ~ '\n\n' }}
+        {%- endif -%}
+        {%- if content -%}
+            {{- content }}
+        {%- endif -%}
+        {%- if message.tool_calls -%}
+            {{- '\n' ~ toolcall_begin_token ~ '\n' }}
+            {%- for tool_call in message.tool_calls -%}
+                {%- if tool_call.function %}
+                    {%- set tool_call = tool_call.function %}
+                {%- endif %}
+                {{- '<invoke name="' + tool_call.name + '">' }}
+                {% set _args = tool_call.arguments %}
+                {%- for k, v in _args.items() %}
+                {{- '<parameter name="' + k + '">' }}
+                {{- v | tojson(ensure_ascii=False) if v is not string else v }}
+                {{- '</parameter>' }}
+                {% endfor %}
+                {{- '</invoke>' ~ '\n' }}
+            {%- endfor -%}
+            {{- toolcall_end_token}}
+            {%- set last_tool_call.name = message.tool_calls[-1].name -%}
+        {%- else -%}
+            {%- set last_tool_call.name = none -%}
+        {%- endif -%}
+        {{- '[e~[' ~ '\n' }}
+    {%- elif message.role == 'tool' -%}
+    {%- if last_tool_call.name is none -%}
+        {{- raise_exception("Message has tool role, but there was no previous assistant message with a tool call!") }}
+    {%- endif -%}
+    {%- if loop.first or (conversation_messages[loop.index0 - 1].role != 'tool') -%}
+        {{- ']~b]tool' }}
+    {%- endif -%}
+    {%- if message.content is string -%}
+        {{- '\n<response>' }}
+        {{- message.content }}
+        {{- '</response>' }}
+    {%- else -%}
+        {%- for tr in message.content -%}
+            {{- '\n<response>' }}
+            {{- tr.output if tr.output is defined else (tr.text if tr.type == 'text' and tr.text is defined else tr) }}
+            {{- '\n</response>' }}
+        {%- endfor -%}
+    {%- endif -%}
+    {%- if loop.last or (conversation_messages[loop.index0 + 1].role != 'tool') -%}
+        {{- '[e~[\n' -}}
+    {%- endif -%}
+    {%- elif message.role == 'user' -%}
+        {{- ']~b]user' ~ '\n' }}
+        {{- visible_text(message.content) }}
+        {{- '[e~[' ~ '\n' }}
+    {%- endif -%}
+{%- endfor -%}
+{#- Generation prompt -#}
+{%- if add_generation_prompt -%}
+{{- ']~b]ai' ~ '\n' ~ '<think>' ~ '\n' }}
+{%- endif -%}

config.json ADDED Viewed

	@@ -0,0 +1,115 @@

+{
+  "model_name": "List-3.0-Ultra-Coder",
+  "architectures": [
+    "MiniMaxM2ForCausalLM"
+  ],
+  "attn_type_list": [
+    1,
+    1,
+    1,
+    1,
+    1,
+    1,
+    1,
+    1,
+    1,
+    1,
+    1,
+    1,
+    1,
+    1,
+    1,
+    1,
+    1,
+    1,
+    1,
+    1,
+    1,
+    1,
+    1,
+    1,
+    1,
+    1,
+    1,
+    1,
+    1,
+    1,
+    1,
+    1,
+    1,
+    1,
+    1,
+    1,
+    1,
+    1,
+    1,
+    1,
+    1,
+    1,
+    1,
+    1,
+    1,
+    1,
+    1,
+    1,
+    1,
+    1,
+    1,
+    1,
+    1,
+    1,
+    1,
+    1,
+    1,
+    1,
+    1,
+    1,
+    1,
+    1
+  ],
+  "auto_map": {
+    "AutoConfig": "configuration_minimax_m2.MiniMaxM2Config",
+    "AutoModelForCausalLM": "modeling_minimax_m2.MiniMaxM2ForCausalLM"
+  },
+  "dtype": "bfloat16",
+  "head_dim": 128,
+  "hidden_act": "silu",
+  "hidden_size": 3072,
+  "intermediate_size": 1536,
+  "max_position_embeddings": 204800,
+  "model_type": "minimax_m2",
+  "mtp_transformer_layers": 1,
+  "num_attention_heads": 48,
+  "num_experts_per_tok": 8,
+  "num_hidden_layers": 62,
+  "num_key_value_heads": 8,
+  "num_local_experts": 256,
+  "num_mtp_modules": 3,
+  "qk_norm_type": "per_layer",
+  "quantization_config": {
+    "activation_scheme": "dynamic",
+    "fmt": "float8_e4m3fn",
+    "quant_method": "fp8",
+    "weight_block_size": [
+      128,
+      128
+    ],
+    "modules_to_not_convert": [
+      "gate",
+      "e_score_correction_bias",
+      "lm_head"
+    ]
+  },
+  "rms_norm_eps": 1e-06,
+  "rope_theta": 5000000,
+  "rotary_dim": 64,
+  "scoring_func": "sigmoid",
+  "shared_intermediate_size": 0,
+  "tie_word_embeddings": false,
+  "transformers_version": "4.46.1",
+  "use_cache": true,
+  "use_mtp": true,
+  "use_qk_norm": true,
+  "use_routing_bias": true,
+  "vocab_size": 200064
+}

docs/sglang_deploy_guide.md ADDED Viewed

	@@ -0,0 +1,112 @@

+# MiniMax M2.7 Model SGLang Deployment Guide
+[English Version](./sglang_deploy_guide.md) | [Chinese Version](./sglang_deploy_guide_cn.md)
+We recommend using [SGLang](https://github.com/sgl-project/sglang) to deploy the [MiniMax-M2.7](https://huggingface.co/MiniMaxAI/MiniMax-M2.7) model. SGLang is a high-performance inference engine with excellent serving throughput, efficient and intelligent memory management, powerful batch request processing capabilities, and deeply optimized underlying performance. We recommend reviewing SGLang's official documentation to check hardware compatibility before deployment.
+## Applicable Models
+This document applies to the following models. You only need to change the model name during deployment.
+- [MiniMaxAI/MiniMax-M2.7](https://huggingface.co/MiniMaxAI/MiniMax-M2.7)
+- [MiniMaxAI/MiniMax-M2.5](https://huggingface.co/MiniMaxAI/MiniMax-M2.5)
+- [MiniMaxAI/MiniMax-M2.1](https://huggingface.co/MiniMaxAI/MiniMax-M2.1)
+- [MiniMaxAI/MiniMax-M2](https://huggingface.co/MiniMaxAI/MiniMax-M2)
+The deployment process is illustrated below using MiniMax-M2.7 as an example.
+## System Requirements
+- OS: Linux
+- Python: 3.9 - 3.12
+- GPU:
+  - compute capability 7.0 or higher
+  - Memory requirements: 220 GB for weights, 240 GB per 1M context tokens
+The following are recommended configurations; actual requirements should be adjusted based on your use case:
+- **96G x4** GPU: Supports a total KV Cache capacity of 400K tokens.
+- **144G x8** GPU: Supports a total KV Cache capacity of up to 3M tokens.
+> **Note**: The values above represent the total aggregate hardware KV Cache capacity. The maximum context length per individual sequence remains **196K** tokens.
+## Deployment with Python
+It is recommended to use a virtual environment (such as **venv**, **conda**, or **uv**) to avoid dependency conflicts.
+We recommend installing SGLang in a fresh Python environment:
+```bash
+uv venv
+source .venv/bin/activate
+uv pip install sglang
+```
+Run the following command to start the SGLang server. SGLang will automatically download and cache the MiniMax-M2.7 model from Hugging Face.
+4-GPU deployment command:
+```bash
+python -m sglang.launch_server \
+    --model-path MiniMaxAI/MiniMax-M2.7 \
+    --tp-size 4 \
+    --tool-call-parser minimax-m2 \
+    --reasoning-parser minimax-append-think \
+    --host 0.0.0.0 \
+    --trust-remote-code \
+    --port 8000 \
+    --mem-fraction-static 0.85
+```
+8-GPU deployment command:
+```bash
+python -m sglang.launch_server \
+    --model-path MiniMaxAI/MiniMax-M2.7 \
+    --tp-size 8 \
+    --ep-size 8 \
+    --tool-call-parser minimax-m2 \
+    --trust-remote-code \
+    --host 0.0.0.0 \
+    --reasoning-parser minimax-append-think \
+    --port 8000 \
+    --mem-fraction-static 0.85
+```
+## Testing Deployment
+After startup, you can test the SGLang OpenAI-compatible API with the following command:
+```bash
+curl http://localhost:8000/v1/chat/completions \
+    -H "Content-Type: application/json" \
+    -d '{
+        "model": "MiniMaxAI/MiniMax-M2.7",
+        "messages": [
+            {"role": "system", "content": [{"type": "text", "text": "You are a helpful assistant."}]},
+            {"role": "user", "content": [{"type": "text", "text": "Who won the world series in 2020?"}]}
+        ]
+    }'
+```
+## Common Issues
+### MiniMax-M2 model is not currently supported
+Please upgrade to the latest stable version, >= v0.5.4.post1.
+## Getting Support
+If you encounter any issues while deploying the MiniMax model:
+- Contact our technical support team through official channels such as email at [model@minimax.io](mailto:model@minimax.io)
+- Submit an issue on our [GitHub](https://github.com/MiniMax-AI) repository
+We continuously optimize the deployment experience for our models. Feedback is welcome!

docs/sglang_deploy_guide_cn.md ADDED Viewed

	@@ -0,0 +1,121 @@

+# MiniMax M2.7 模型 SGLang 部署指南
+[英文版](./sglang_deploy_guide.md) | [中文版](./sglang_deploy_guide_cn.md)
+我们推荐使用 [SGLang](https://github.com/sgl-project/sglang) 来部署 [MiniMax-M2.7](https://huggingface.co/MiniMaxAI/MiniMax-M2.7) 模型。SGLang 是一个高性能的推理引擎，其具有卓越的服务吞吐、高效智能的内存管理机制、强大的批量请求处理能力、深度优化的底层性能等特性。我们建议在部署之前查看 SGLang 的官方文档以检查硬件兼容性。
+## 本文档适用模型
+本文档适用以下模型，只需在部署时修改模型名称即可。
+- [MiniMaxAI/MiniMax-M2.7](https://huggingface.co/MiniMaxAI/MiniMax-M2.7)
+- [MiniMaxAI/MiniMax-M2.5](https://huggingface.co/MiniMaxAI/MiniMax-M2.5)
+- [MiniMaxAI/MiniMax-M2.1](https://huggingface.co/MiniMaxAI/MiniMax-M2.1)
+- [MiniMaxAI/MiniMax-M2](https://huggingface.co/MiniMaxAI/MiniMax-M2)
+以下以 MiniMax-M2.7 为例说明部署流程。
+## 环境要求
+- OS：Linux
+- Python：3.9 - 3.12
+- GPU：
+  - compute capability 7.0 or higher
+  - 显存需求：权重需要 220 GB，每 1M 上下文 token 需要 240 GB
+以下为推荐配置，实际需求请根据业务场景调整：
+- **96G x4 GPU**：总 KV Cache 容量支持 40 万 token。
+- **144G x8 GPU**：总 KV Cache 容量支持高达 300 万 token。
+> **注**：以上数值为硬件支持的最大并发缓存总量，模型单序列（Single Sequence）长度上限仍为 196k。
+## 使用 Python 部署
+建议使用虚拟环境（如 **venv**、**conda**、**uv**）以避免依赖冲突。
+建议在全新的 Python 环境中安装 SGLang:
+```bash
+uv venv
+source .venv/bin/activate
+uv pip install sglang
+```
+运行如下命令启动 SGLang 服务器，SGLang 会自动从 Huggingface 下载并缓存 MiniMax-M2.7 模型。
+4 卡部署命令：
+```bash
+python -m sglang.launch_server \
+    --model-path MiniMaxAI/MiniMax-M2.7 \
+    --tp-size 4 \
+    --tool-call-parser minimax-m2 \
+    --reasoning-parser minimax-append-think \
+    --host 0.0.0.0 \
+    --trust-remote-code \
+    --port 8000 \
+    --mem-fraction-static 0.85
+```
+8 卡部署命令：
+```bash
+python -m sglang.launch_server \
+    --model-path MiniMaxAI/MiniMax-M2.7 \
+    --tp-size 8 \
+    --ep-size 8 \
+    --tool-call-parser minimax-m2 \
+    --trust-remote-code \
+    --host 0.0.0.0 \
+    --reasoning-parser minimax-append-think \
+    --port 8000 \
+    --mem-fraction-static 0.85
+```
+## 测试部署
+启动后，可以通过如下命令测试 SGLang OpenAI 兼容接口：
+```bash
+curl http://localhost:8000/v1/chat/completions \
+    -H "Content-Type: application/json" \
+    -d '{
+        "model": "MiniMaxAI/MiniMax-M2.7",
+        "messages": [
+            {"role": "system", "content": [{"type": "text", "text": "You are a helpful assistant."}]},
+            {"role": "user", "content": [{"type": "text", "text": "Who won the world series in 2020?"}]}
+        ]
+    }'
+```
+## 常见问题
+### Huggingface 网络问题
+如果遇到网络问题，可以设置代理后再进行拉取。
+```bash
+export HF_ENDPOINT=https://hf-mirror.com
+```
+### MiniMax-M2 model is not currently supported
+请升级到最新的稳定版本, >= v0.5.4.post1.
+## 获取支持
+如果在部署 MiniMax 模型过程中遇到任何问题：
+- 通过邮箱 [model@minimax.io](mailto:model@minimax.io) 等官方渠道联系我们的技术支持团队
+- 在我们的 [GitHub](https://github.com/MiniMax-AI) 仓库提交 Issue
+- 通过我们的 [官方企业微信交流群](https://github.com/MiniMax-AI/MiniMax-AI.github.io/blob/main/images/wechat-qrcode.jpeg) 反馈
+我们会持续优化模型的部署体验，欢迎反馈！

docs/tool_calling_guide.md ADDED Viewed

	@@ -0,0 +1,487 @@

+# MiniMax-M2.7 Tool Calling Guide
+[English Version](./tool_calling_guide.md) | [Chinese Version](./tool_calling_guide_cn.md)
+MiniMax-M2.7 supports the same toolcall syntax as MiniMax-M2.
+## Introduction
+The MiniMax-M2.7 model supports tool calling capabilities, enabling the model to identify when external tools need to be called and output tool call parameters in a structured format. This document provides detailed instructions on how to use the tool calling features of MiniMax-M2.7.
+## Basic Example
+The following Python script implements a weather query tool call example based on the OpenAI SDK:
+```python
+from openai import OpenAI
+import json
+client = OpenAI(base_url="http://localhost:8000/v1", api_key="dummy")
+def get_weather(location: str, unit: str):
+    return f"Getting the weather for {location} in {unit}..."
+tool_functions = {"get_weather": get_weather}
+tools = [{
+    "type": "function",
+    "function": {
+        "name": "get_weather",
+        "description": "Get the current weather in a given location",
+        "parameters": {
+            "type": "object",
+            "properties": {
+                "location": {"type": "string", "description": "City and state, e.g., 'San Francisco, CA'"},
+                "unit": {"type": "string", "enum": ["celsius", "fahrenheit"]}
+            },
+            "required": ["location", "unit"]
+        }
+    }
+}]
+response = client.chat.completions.create(
+    model=client.models.list().data[0].id,
+    messages=[{"role": "user", "content": "What's the weather like in San Francisco? use celsius."}],
+    tools=tools,
+    tool_choice="auto"
+)
+print(response)
+tool_call = response.choices[0].message.tool_calls[0].function
+print(f"Function called: {tool_call.name}")
+print(f"Arguments: {tool_call.arguments}")
+print(f"Result: {get_weather(**json.loads(tool_call.arguments))}")
+```
+**Output Example:**
+```
+Function called: get_weather
+Arguments: {"location": "San Francisco, CA", "unit": "celsius"}
+Result: Getting the weather for San Francisco, CA in celsius...
+```
+## Manually Parsing Model Output
+**We strongly recommend using vLLM or SGLang for parsing tool calls.** If you cannot use the built-in parser of inference engines (e.g., vLLM and SGLang) that support MiniMax-M2.7, or need to use other inference frameworks (such as transformers, TGI, etc.), you can manually parse the model's raw output using the following method. This approach requires you to parse the XML tag format of the model output yourself.
+### Example Using Transformers
+Here is a complete example using the transformers library:
+```python
+from transformers import AutoTokenizer
+def get_default_tools():
+    return [
+        {
+          "name": "get_current_weather",
+          "description": "Get the latest weather for a location",
+          "parameters": {
+              "type": "object",
+              "properties": {
+                  "location": {
+                      "type": "string",
+                      "description": "A certain city, such as Beijing, Shanghai"
+                  }
+              },
+          }
+          "required": ["location"],
+          "type": "object"
+        }
+    ]
+# Load model and tokenizer
+tokenizer = AutoTokenizer.from_pretrained(model_id)
+prompt = "What's the weather like in Shanghai today?"
+messages = [
+    {"role": "system", "content": "You are a helpful assistant."},
+    {"role": "user", "content": prompt},
+]
+# Enable function calling tools
+tools = get_default_tools()
+# Apply chat template and include tool definitions
+text = tokenizer.apply_chat_template(
+    messages,
+    tokenize=False,
+    add_generation_prompt=True,
+    tools=tools
+)
+# Send request (using any inference service)
+import requests
+payload = {
+    "model": "MiniMaxAI/MiniMax-M2.7",
+    "prompt": text,
+    "max_tokens": 4096
+}
+response = requests.post(
+    "http://localhost:8000/v1/completions",
+    headers={"Content-Type": "application/json"},
+    json=payload,
+    stream=False,
+)
+# Model output needs manual parsing
+raw_output = response.json()["choices"][0]["text"]
+print("Raw output:", raw_output)
+# Use the parsing function below to process the output
+tool_calls = parse_tool_calls(raw_output, tools)
+```
+## 🛠️ Tool Call Definition
+### Tool Structure
+Tool calls need to define the `tools` field in the request body. Each tool consists of the following parts:
+```json
+{
+  "tools": [
+    {
+      "name": "search_web",
+      "description": "Search function.",
+      "parameters": {
+        "properties": {
+          "query_list": {
+            "description": "Keywords for search, list should contain 1 element.",
+            "items": { "type": "string" },
+            "type": "array"
+          },
+          "query_tag": {
+            "description": "Category of query",
+            "items": { "type": "string" },
+            "type": "array"
+          }
+        },
+        "required": [ "query_list", "query_tag" ],
+        "type": "object"
+      }
+    }
+  ]
+}
+```
+**Field Descriptions:**
+- `name`: Function name
+- `description`: Function description
+- `parameters`: Function parameter definition
+  - `properties`: Parameter property definition, where key is the parameter name and value contains detailed parameter description
+  - `required`: List of required parameters
+  - `type`: Parameter type (usually "object")
+### Internal Processing Format
+When processing within the MiniMax-M2.7 model, tool definitions are converted to a special format and concatenated to the input text. Here is a complete example:
+```
+]~!b[]~b]system
+You are a helpful assistant.
+# Tools
+You may call one or more tools to assist with the user query.
+Here are the tools available in JSONSchema format:
+<tools>
+<tool>{"name": "search_web", "description": "Search function.", "parameters": {"type": "object", "properties": {"query_list": {"type": "array", "items": {"type": "string"}, "description": "Keywords for search, list should contain 1 element."}, "query_tag": {"type": "array", "items": {"type": "string"}, "description": "Category of query"}}, "required": ["query_list", "query_tag"]}}</tool>
+</tools>
+When making tool calls, use XML format to invoke tools and pass parameters:
+<minimax:tool_call>
+<invoke name="tool-name-1">
+<parameter name="param-key-1">param-value-1</parameter>
+<parameter name="param-key-2">param-value-2</parameter>
+...
+</invoke>
+[e~[
+]~b]user
+When were the latest announcements from OpenAI and Gemini?[e~[
+]~b]ai
+<think>
+```
+**Format Description:**
+- `]~!b[]~b]system`: System message start marker
+- `[e~[`: Message end marker
+- `]~b]user`: User message start marker
+- `]~b]ai`: Assistant message start marker
+- `]~b]tool`: Tool result message start marker
+- `<tools>...</tools>`: Tool definition area, each tool is wrapped with `<tool>` tag, content is JSON Schema
+- `<minimax:tool_call>...</minimax:tool_call>`: Tool call area
+- `<think>...</think>`: Thinking process marker during generation
+### Model Output Format
+MiniMax-M2.7 uses structured XML tag format:
+```xml
+<minimax:tool_call>
+<invoke name="search_web">
+<parameter name="query_tag">["technology", "events"]</parameter>
+<parameter name="query_list">["\"OpenAI\" \"latest\" \"release\""]</parameter>
+</invoke>
+<invoke name="search_web">
+<parameter name="query_tag">["technology", "events"]</parameter>
+<parameter name="query_list">["\"Gemini\" \"latest\" \"release\""]</parameter>
+</invoke>
+</minimax:tool_call>
+```
+Each tool call uses the `<invoke name="function_name">` tag, and parameters use the `<parameter name="parameter_name">` tag wrapper.
+## Manually Parsing Tool Call Results
+### Parsing Tool Calls
+MiniMax-M2.7 uses structured XML tags, which require a different parsing approach. The core function is as follows:
+```python
+import re
+import json
+from typing import Any, Optional, List, Dict
+def extract_name(name_str: str) -> str:
+    """Extract name from quoted string"""
+    name_str = name_str.strip()
+    if name_str.startswith('"') and name_str.endswith('"'):
+        return name_str[1:-1]
+    elif name_str.startswith("'") and name_str.endswith("'"):
+        return name_str[1:-1]
+    return name_str
+def convert_param_value(value: str, param_type: str) -> Any:
+    """Convert parameter value based on parameter type"""
+    if value.lower() == "null":
+        return None
+    param_type = param_type.lower()
+    if param_type in ["string", "str", "text"]:
+        return value
+    elif param_type in ["integer", "int"]:
+        try:
+            return int(value)
+        except (ValueError, TypeError):
+            return value
+    elif param_type in ["number", "float"]:
+        try:
+            val = float(value)
+            return val if val != int(val) else int(val)
+        except (ValueError, TypeError):
+            return value
+    elif param_type in ["boolean", "bool"]:
+        return value.lower() in ["true", "1"]
+    elif param_type in ["object", "array"]:
+        try:
+            return json.loads(value)
+        except json.JSONDecodeError:
+            return value
+    else:
+        # Try JSON parsing, return string if failed
+        try:
+            return json.loads(value)
+        except json.JSONDecodeError:
+            return value
+def parse_tool_calls(model_output: str, tools: Optional[List[Dict]] = None) -> List[Dict]:
+    """
+    Extract all tool calls from model output
+    Args:
+        model_output: Complete output text from the model
+        tools: Tool definition list for getting parameter type information, format can be:
+               - [{"name": "...", "parameters": {...}}]
+               - [{"type": "function", "function": {"name": "...", "parameters": {...}}}]
+    Returns:
+        Parsed tool call list, each element contains name and arguments fields
+    Example:
+        >>> tools = [{
+        ...     "name": "get_weather",
+        ...     "parameters": {
+        ...         "type": "object",
+        ...         "properties": {
+        ...             "location": {"type": "string"},
+        ...             "unit": {"type": "string"}
+        ...         }
+        ...     }
+        ... }]
+        >>> output = '''<minimax:tool_call>
+        ... <invoke name="get_weather">
+        ... <parameter name="location">San Francisco</parameter>
+        ... <parameter name="unit">celsius</parameter>
+        ... </invoke>
+        ... </minimax:tool_call>'''
+        >>> result = parse_tool_calls(output, tools)
+        >>> print(result)
+        [{'name': 'get_weather', 'arguments': {'location': 'San Francisco', 'unit': 'celsius'}}]
+    """
+    # Quick check if tool call marker is present
+    if "<minimax:tool_call>" not in model_output:
+        return []
+    tool_calls = []
+    try:
+        # Match all <minimax:tool_call> blocks
+        tool_call_regex = re.compile(r"<minimax:tool_call>(.*?)</minimax:tool_call>", re.DOTALL)
+        invoke_regex = re.compile(r"<invoke name=(.*?)</invoke>", re.DOTALL)
+        parameter_regex = re.compile(r"<parameter name=(.*?)</parameter>", re.DOTALL)
+        # Iterate through all tool_call blocks
+        for tool_call_match in tool_call_regex.findall(model_output):
+            # Iterate through all invokes in this block
+            for invoke_match in invoke_regex.findall(tool_call_match):
+                # Extract function name
+                name_match = re.search(r'^([^>]+)', invoke_match)
+                if not name_match:
+                    continue
+                function_name = extract_name(name_match.group(1))
+                # Get parameter configuration
+                param_config = {}
+                if tools:
+                    for tool in tools:
+                        tool_name = tool.get("name") or tool.get("function", {}).get("name")
+                        if tool_name == function_name:
+                            params = tool.get("parameters") or tool.get("function", {}).get("parameters")
+                            if isinstance(params, dict) and "properties" in params:
+                                param_config = params["properties"]
+                            break
+                # Extract parameters
+                param_dict = {}
+                for match in parameter_regex.findall(invoke_match):
+                    param_match = re.search(r'^([^>]+)>(.*)', match, re.DOTALL)
+                    if param_match:
+                        param_name = extract_name(param_match.group(1))
+                        param_value = param_match.group(2).strip()
+                        # Remove leading and trailing newlines
+                        if param_value.startswith('\n'):
+                            param_value = param_value[1:]
+                        if param_value.endswith('\n'):
+                            param_value = param_value[:-1]
+                        # Get parameter type and convert
+                        param_type = "string"
+                        if param_name in param_config:
+                            if isinstance(param_config[param_name], dict) and "type" in param_config[param_name]:
+                                param_type = param_config[param_name]["type"]
+                        param_dict[param_name] = convert_param_value(param_value, param_type)
+                tool_calls.append({
+                    "name": function_name,
+                    "arguments": param_dict
+                })
+    except Exception as e:
+        print(f"Failed to parse tool calls: {e}")
+        return []
+    return tool_calls
+```
+**Usage Example:**
+```python
+# Define tools
+tools = [
+    {
+        "name": "get_weather",
+        "parameters": {
+            "type": "object",
+            "properties": {
+                "location": {"type": "string"},
+                "unit": {"type": "string"}
+            },
+            "required": ["location", "unit"]
+        }
+    }
+]
+# Model output
+model_output = """Let me help you query the weather.
+<minimax:tool_call>
+<invoke name="get_weather">
+<parameter name="location">San Francisco</parameter>
+<parameter name="unit">celsius</parameter>
+</invoke>
+</minimax:tool_call>"""
+# Parse tool calls
+tool_calls = parse_tool_calls(model_output, tools)
+# Output results
+for call in tool_calls:
+    print(f"Function called: {call['name']}")
+    print(f"Arguments: {call['arguments']}")
+    # Output: Function called: get_weather
+    #         Arguments: {'location': 'San Francisco', 'unit': 'celsius'}
+```
+### Executing Tool Calls
+After parsing is complete, you can execute the corresponding tool and construct the return result:
+```python
+def execute_function_call(function_name: str, arguments: dict):
+    """Execute function call and return result"""
+    if function_name == "get_weather":
+        location = arguments.get("location", "Unknown location")
+        unit = arguments.get("unit", "celsius")
+        # Build function execution result
+        return {
+            "role": "tool",
+            "content": [
+              {
+                "name": function_name,
+                "type": "text",
+                "text": json.dumps({
+                    "location": location,
+                    "temperature": "25",
+                    "unit": unit,
+                    "weather": "Sunny"
+                }, ensure_ascii=False)
+              }
+            ]
+          }
+    elif function_name == "search_web":
+        query_list = arguments.get("query_list", [])
+        query_tag = arguments.get("query_tag", [])
+        # Simulate search results
+        return {
+            "role": "tool",
+            "content": [
+              {
+                "name": function_name,
+                "type": "text",
+                "text": f"Search keywords: {query_list}, Category: {query_tag}\nSearch results: Relevant information found"
+              }
+            ]
+          }
+    return None
+```
+### Returning Tool Execution Results to the Model
+After successfully parsing tool calls, you should add the tool execution results to the conversation history so that the model can access and utilize this information in subsequent interactions. Refer to [chat_template.jinja](https://huggingface.co/MiniMaxAI/MiniMax-M2.7/blob/main/chat_template.jinja) for concatenation format.
+## References
+- [MiniMax-M2.7 Model Repository](https://github.com/MiniMax-AI/MiniMax-M2.7)
+- [vLLM Project Homepage](https://github.com/vllm-project/vllm)
+- [SGLang Project Homepage](https://github.com/sgl-project/sglang)
+- [OpenAI Python SDK](https://github.com/openai/openai-python)

docs/tool_calling_guide_cn.md ADDED Viewed

	@@ -0,0 +1,499 @@

+# MiniMax-M2.7 工具调用指南
+[英文版](./tool_calling_guide.md) | [中文版](./tool_calling_guide_cn.md)
+MiniMax-M2.7 支持与 MiniMax-M2 相同的工具调用语法。
+## 简介
+MiniMax-M2.7 模型支持工具调用功能，使模型能够识别何时需要调用外部工具，并以结构化格式输出工具调用参数。本文档提供了有关如何使用 MiniMax-M2.7 工具调用功能的详细说明。
+## 基础示例
+以下 Python 脚本基于 OpenAI SDK 实现了一个天气查询工具调用示例：
+```python
+from openai import OpenAI
+import json
+client = OpenAI(base_url="http://localhost:8000/v1", api_key="dummy")
+def get_weather(location: str, unit: str):
+    return f"Getting the weather for {location} in {unit}..."
+tool_functions = {"get_weather": get_weather}
+tools = [{
+    "type": "function",
+    "function": {
+        "name": "get_weather",
+        "description": "Get the current weather in a given location",
+        "parameters": {
+            "type": "object",
+            "properties": {
+                "location": {"type": "string", "description": "City and state, e.g., 'San Francisco, CA'"},
+                "unit": {"type": "string", "enum": ["celsius", "fahrenheit"]}
+            },
+            "required": ["location", "unit"]
+        }
+    }
+}]
+response = client.chat.completions.create(
+    model=client.models.list().data[0].id,
+    messages=[{"role": "user", "content": "What's the weather like in San Francisco? use celsius."}],
+    tools=tools,
+    tool_choice="auto"
+)
+print(response)
+tool_call = response.choices[0].message.tool_calls[0].function
+print(f"Function called: {tool_call.name}")
+print(f"Arguments: {tool_call.arguments}")
+print(f"Result: {get_weather(**json.loads(tool_call.arguments))}")
+```
+**输出示例：**
+```
+Function called: get_weather
+Arguments: {"location": "San Francisco, CA", "unit": "celsius"}
+Result: Getting the weather for San Francisco, CA in celsius...
+```
+## 手动解析模型输出
+**我们强烈建议使用 vLLM 或 SGLnag 来解析工具调用。** 如果您无法使用支持 MiniMax-M2.7 的推理引擎（如 vLLM 和 SGLang）的内置解析器，或需要使用其他推理框架（如 transformers、TGI 等），您可以使用以下方法手动解析模型的原始输出。这种方法需要您自己解析模型输出的 XML 标签格式。
+### 使用 Transformers 的示例
+这是一个使用 transformers 库的完整示例：
+```python
+from transformers import AutoTokenizer
+def get_default_tools():
+    return [
+        {
+          "name": "get_current_weather",
+          "description": "Get the latest weather for a location",
+          "parameters": {
+              "type": "object",
+              "properties": {
+                  "location": {
+                      "type": "string",
+                      "description": "A certain city, such as Beijing, Shanghai"
+                  }
+              },
+          }
+          "required": ["location"],
+          "type": "object"
+        }
+    ]
+# Load model and tokenizer
+tokenizer = AutoTokenizer.from_pretrained(model_id)
+prompt = "What's the weather like in Shanghai today?"
+messages = [
+    {"role": "system", "content": "You are a helpful assistant."},
+    {"role": "user", "content": prompt},
+]
+# Enable function calling tools
+tools = get_default_tools()
+# Apply chat template and include tool definitions
+text = tokenizer.apply_chat_template(
+    messages,
+    tokenize=False,
+    add_generation_prompt=True,
+    tools=tools
+)
+# Send request (using any inference service)
+import requests
+payload = {
+    "model": "MiniMaxAI/MiniMax-M2.7",
+    "prompt": text,
+    "max_tokens": 4096
+}
+response = requests.post(
+    "http://localhost:8000/v1/completions",
+    headers={"Content-Type": "application/json"},
+    json=payload,
+    stream=False,
+)
+# Model output needs manual parsing
+raw_output = response.json()["choices"][0]["text"]
+print("Raw output:", raw_output)
+# Use the parsing function below to process the output
+tool_calls = parse_tool_calls(raw_output, tools)
+```
+## 🛠️ 工具调用定义
+### 工具结构
+工具调用需要在请求体中定义 `tools` 字段。每个工具由以下部分组成：
+```json
+{
+  "tools": [
+    {
+      "name": "search_web",
+      "description": "Search function.",
+      "parameters": {
+        "properties": {
+          "query_list": {
+            "description": "Keywords for search, list should contain 1 element.",
+            "items": { "type": "string" },
+            "type": "array"
+          },
+          "query_tag": {
+            "description": "Category of query",
+            "items": { "type": "string" },
+            "type": "array"
+          }
+        },
+        "required": [ "query_list", "query_tag" ],
+        "type": "object"
+      }
+    }
+  ]
+}
+```
+**字段说明：**
+- `name`：函数名称
+- `description`：函数描述
+- `parameters`：函数参数定义
+  - `properties`：参数属性定义，其中键是参数名称，值包含详细的参数描述
+  - `required`：必需参数列表
+  - `type`：参数类型（通常为 "object"）
+### 内部处理格式
+在 MiniMax-M2.7 模型内部处理时，工具定义会被转换为特殊格式并连接到输入文本中。以下是一个完整示例：
+```
+]~!b[]~b]system
+You are a helpful assistant.
+# Tools
+You may call one or more tools to assist with the user query.
+Here are the tools available in JSONSchema format:
+<tools>
+<tool>{"name": "search_web", "description": "Search function.", "parameters": {"type": "object", "properties": {"query_list": {"type": "array", "items": {"type": "string"}, "description": "Keywords for search, list should contain 1 element."}, "query_tag": {"type": "array", "items": {"type": "string"}, "description": "Category of query"}}, "required": ["query_list", "query_tag"]}}</tool>
+</tools>
+When making tool calls, use XML format to invoke tools and pass parameters:
+<minimax:tool_call>
+<invoke name="tool-name-1">
+<parameter name="param-key-1">param-value-1</parameter>
+<parameter name="param-key-2">param-value-2</parameter>
+...
+</invoke>
+[e~[
+]~b]user
+When were the latest announcements from OpenAI and Gemini?[e~[
+]~b]ai
+<think>
+```
+**格式说明：**
+- `]~!b[]~b]system`：系统消息开始标记
+- `[e~[`：消息结束标记
+- `]~b]user`：用户消息开始标记
+- `]~b]ai`：助手消息开始标记
+- `]~b]tool`：工具结果消息开始标记
+- `<tools>...</tools>`：工具定义区域，每个工具都用 `<tool>` 标签包装，内容为 JSON Schema
+- `<minimax:tool_call>...</minimax:tool_call>`：工具调用区域
+- `<think>...</think>`：生成过程中的思考过程标记
+### 模型输出格式
+MiniMax-M2.7 使用结构化的 XML 标签格式：
+```xml
+<minimax:tool_call>
+<invoke name="search_web">
+<parameter name="query_tag">["technology", "events"]</parameter>
+<parameter name="query_list">["\"OpenAI\" \"latest\" \"release\""]</parameter>
+</invoke>
+<invoke name="search_web">
+<parameter name="query_tag">["technology", "events"]</parameter>
+<parameter name="query_list">["\"Gemini\" \"latest\" \"release\""]</parameter>
+</invoke>
+</minimax:tool_call>
+```
+每个工具调用使用 `<invoke name="function_name">` 标签，参数使用 `<parameter name="parameter_name">` 标签包装。
+## 手动解析工具调用结果
+### 解析工具调用
+MiniMax-M2.7 使用结构化的 XML 标签，这需要一种不同的解析方法。核心函数如下：
+```python
+import re
+import json
+from typing import Any, Optional, List, Dict
+def extract_name(name_str: str) -> str:
+    """Extract name from quoted string"""
+    name_str = name_str.strip()
+    if name_str.startswith('"') and name_str.endswith('"'):
+        return name_str[1:-1]
+    elif name_str.startswith("'") and name_str.endswith("'"):
+        return name_str[1:-1]
+    return name_str
+def convert_param_value(value: str, param_type: str) -> Any:
+    """Convert parameter value based on parameter type"""
+    if value.lower() == "null":
+        return None
+    param_type = param_type.lower()
+    if param_type in ["string", "str", "text"]:
+        return value
+    elif param_type in ["integer", "int"]:
+        try:
+            return int(value)
+        except (ValueError, TypeError):
+            return value
+    elif param_type in ["number", "float"]:
+        try:
+            val = float(value)
+            return val if val != int(val) else int(val)
+        except (ValueError, TypeError):
+            return value
+    elif param_type in ["boolean", "bool"]:
+        return value.lower() in ["true", "1"]
+    elif param_type in ["object", "array"]:
+        try:
+            return json.loads(value)
+        except json.JSONDecodeError:
+            return value
+    else:
+        # Try JSON parsing, return string if failed
+        try:
+            return json.loads(value)
+        except json.JSONDecodeError:
+            return value
+def parse_tool_calls(model_output: str, tools: Optional[List[Dict]] = None) -> List[Dict]:
+    """
+    Extract all tool calls from model output
+    Args:
+        model_output: Complete output text from the model
+        tools: Tool definition list for getting parameter type information, format can be:
+               - [{"name": "...", "parameters": {...}}]
+               - [{"type": "function", "function": {"name": "...", "parameters": {...}}}]
+    Returns:
+        Parsed tool call list, each element contains name and arguments fields
+    Example:
+        >>> tools = [{
+        ...     "name": "get_weather",
+        ...     "parameters": {
+        ...         "type": "object",
+        ...         "properties": {
+        ...             "location": {"type": "string"},
+        ...             "unit": {"type": "string"}
+        ...         }
+        ...     }
+        ... }]
+        >>> output = '''<minimax:tool_call>
+        ... <invoke name="get_weather">
+        ... <parameter name="location">San Francisco</parameter>
+        ... <parameter name="unit">celsius</parameter>
+        ... </invoke>
+        ... </minimax:tool_call>'''
+        >>> result = parse_tool_calls(output, tools)
+        >>> print(result)
+        [{'name': 'get_weather', 'arguments': {'location': 'San Francisco', 'unit': 'celsius'}}]
+    """
+    # Quick check if tool call marker is present
+    if "<minimax:tool_call>" not in model_output:
+        return []
+    tool_calls = []
+    try:
+        # Match all <minimax:tool_call> blocks
+        tool_call_regex = re.compile(r"<minimax:tool_call>(.*?)</minimax:tool_call>", re.DOTALL)
+        invoke_regex = re.compile(r"<invoke name=(.*?)</invoke>", re.DOTALL)
+        parameter_regex = re.compile(r"<parameter name=(.*?)</parameter>", re.DOTALL)
+        # Iterate through all tool_call blocks
+        for tool_call_match in tool_call_regex.findall(model_output):
+            # Iterate through all invokes in this block
+            for invoke_match in invoke_regex.findall(tool_call_match):
+                # Extract function name
+                name_match = re.search(r'^([^>]+)', invoke_match)
+                if not name_match:
+                    continue
+                function_name = extract_name(name_match.group(1))
+                # Get parameter configuration
+                param_config = {}
+                if tools:
+                    for tool in tools:
+                        tool_name = tool.get("name") or tool.get("function", {}).get("name")
+                        if tool_name == function_name:
+                            params = tool.get("parameters") or tool.get("function", {}).get("parameters")
+                            if isinstance(params, dict) and "properties" in params:
+                                param_config = params["properties"]
+                            break
+                # Extract parameters
+                param_dict = {}
+                for match in parameter_regex.findall(invoke_match):
+                    param_match = re.search(r'^([^>]+)>(.*)', match, re.DOTALL)
+                    if param_match:
+                        param_name = extract_name(param_match.group(1))
+                        param_value = param_match.group(2).strip()
+                        # Remove leading and trailing newlines
+                        if param_value.startswith('\n'):
+                            param_value = param_value[1:]
+                        if param_value.endswith('\n'):
+                            param_value = param_value[:-1]
+                        # Get parameter type and convert
+                        param_type = "string"
+                        if param_name in param_config:
+                            if isinstance(param_config[param_name], dict) and "type" in param_config[param_name]:
+                                param_type = param_config[param_name]["type"]
+                        param_dict[param_name] = convert_param_value(param_value, param_type)
+                tool_calls.append({
+                    "name": function_name,
+                    "arguments": param_dict
+                })
+    except Exception as e:
+        print(f"Failed to parse tool calls: {e}")
+        return []
+    return tool_calls
+```
+**使用示例：**
+```python
+# Define tools
+tools = [
+    {
+        "name": "get_weather",
+        "parameters": {
+            "type": "object",
+            "properties": {
+                "location": {"type": "string"},
+                "unit": {"type": "string"}
+            },
+            "required": ["location", "unit"]
+        }
+    }
+]
+# Model output
+model_output = """Let me help you query the weather.
+<minimax:tool_call>
+<invoke name="get_weather">
+<parameter name="location">San Francisco</parameter>
+<parameter name="unit">celsius</parameter>
+</invoke>
+</minimax:tool_call>"""
+# Parse tool calls
+tool_calls = parse_tool_calls(model_output, tools)
+# Output results
+for call in tool_calls:
+    print(f"Function called: {call['name']}")
+    print(f"Arguments: {call['arguments']}")
+    # Output: Function called: get_weather
+    #         Arguments: {'location': 'San Francisco', 'unit': 'celsius'}
+```
+### 执行工具调用
+完成解析后，您可以执行相应的工具并构造返回结果：
+```python
+def execute_function_call(function_name: str, arguments: dict):
+    """Execute function call and return result"""
+    if function_name == "get_weather":
+        location = arguments.get("location", "Unknown location")
+        unit = arguments.get("unit", "celsius")
+        # Build function execution result
+        return {
+            "role": "tool",
+            "content": [
+              {
+                "name": function_name,
+                "type": "text",
+                "text": json.dumps({
+                    "location": location,
+                    "temperature": "25",
+                    "unit": unit,
+                    "weather": "Sunny"
+                }, ensure_ascii=False)
+              }
+            ]
+          }
+    elif function_name == "search_web":
+        query_list = arguments.get("query_list", [])
+        query_tag = arguments.get("query_tag", [])
+        # Simulate search results
+        return {
+            "role": "tool",
+            "content": [
+              {
+                "name": function_name,
+                "type": "text",
+                "text": f"Search keywords: {query_list}, Category: {query_tag}\nSearch results: Relevant information found"
+              }
+            ]
+          }
+    return None
+```
+### 将工具执行结果返回给模型
+在成功解析工具调用后，您应该将工具执行结果添加到对话历史中，以便模型在后续交互中可以访问和利用这些信息。请参考 [chat_template.jinja](https://huggingface.co/MiniMaxAI/MiniMax-M2.7/blob/main/chat_template.jinja) 了解连接格式。
+## 参考文献
+- [MiniMax-M2.7 模型仓库](https://github.com/MiniMax-AI/MiniMax-M2.7)
+- [vLLM 项目主页](https://github.com/vllm-project/vllm)
+- [SGLang 项目主页](https://github.com/sgl-project/sglang)
+- [OpenAI Python SDK](https://github.com/openai/openai-python)
+## 获取支持
+如果遇到任何问题：
+- 通过邮箱 [model@minimax.io](mailto:model@minimax.io) 等官方渠道联系我们的技术支持团队
+- 在我们的仓库提交 Issue
+- 通过我们的 [官方企业微信交流群](https://github.com/MiniMax-AI/MiniMax-AI.github.io/blob/main/images/wechat-qrcode.jpeg) 反馈
+我们会持续优化模型的使用体验，欢迎反馈！

docs/transformers_deploy_guide.md ADDED Viewed

	@@ -0,0 +1,93 @@

+# MiniMax M2.7 Model Transformers Deployment Guide
+[English Version](./transformers_deploy_guide.md) | [Chinese Version](./transformers_deploy_guide_cn.md)
+## Applicable Models
+This document applies to the following models. You only need to change the model name during deployment.
+- [MiniMaxAI/MiniMax-M2.7](https://huggingface.co/MiniMaxAI/MiniMax-M2.7)
+- [MiniMaxAI/MiniMax-M2.5](https://huggingface.co/MiniMaxAI/MiniMax-M2.5)
+- [MiniMaxAI/MiniMax-M2.1](https://huggingface.co/MiniMaxAI/MiniMax-M2.1)
+- [MiniMaxAI/MiniMax-M2](https://huggingface.co/MiniMaxAI/MiniMax-M2)
+The deployment process is illustrated below using MiniMax-M2.7 as an example.
+## System Requirements
+- OS: Linux
+- Python: 3.9 - 3.12
+- Transformers: 4.57.1
+- GPU:
+  - compute capability 7.0 or higher
+  - Memory requirements: 220 GB for weights.
+## Deployment with Python
+It is recommended to use a virtual environment (such as **venv**, **conda**, or **uv**) to avoid dependency conflicts.
+We recommend installing Transformers in a fresh Python environment:
+```bash
+uv pip install transformers==4.57.1 torch accelerate --torch-backend=auto
+```
+Run the following Python script to run the model. Transformers will automatically download and cache the MiniMax-M2.7 model from Hugging Face.
+```python
+from transformers import AutoModelForCausalLM, AutoTokenizer, GenerationConfig
+import torch
+MODEL_PATH = "MiniMaxAI/MiniMax-M2.7"
+model = AutoModelForCausalLM.from_pretrained(
+    MODEL_PATH,
+    device_map="auto",
+    trust_remote_code=True,
+)
+tokenizer = AutoTokenizer.from_pretrained(MODEL_PATH)
+messages = [
+    {"role": "user", "content": [{"type": "text", "text": "What is your favourite condiment?"}]},
+    {"role": "assistant", "content": [{"type": "text", "text": "Well, I'm quite partial to a good squeeze of fresh lemon juice. It adds just the right amount of zesty flavour to whatever I'm cooking up in the kitchen!"}]},
+    {"role": "user", "content": [{"type": "text", "text": "Do you have mayonnaise recipes?"}]}
+]
+model_inputs = tokenizer.apply_chat_template(messages, return_tensors="pt", add_generation_prompt=True).to("cuda")
+generated_ids = model.generate(model_inputs, max_new_tokens=100, generation_config=model.generation_config)
+response = tokenizer.batch_decode(generated_ids)[0]
+print(response)
+```
+## Common Issues
+### Hugging Face Network Issues
+If you encounter network issues, you can set up a proxy before pulling the model.
+```bash
+export HF_ENDPOINT=https://hf-mirror.com
+```
+### MiniMax-M2 model is not currently supported
+Please check that trust_remote_code=True.
+## Getting Support
+If you encounter any issues while deploying the MiniMax model:
+- Contact our technical support team through official channels such as email at [model@minimax.io](mailto:model@minimax.io)
+- Submit an issue on our [GitHub](https://github.com/MiniMax-AI) repository
+We continuously optimize the deployment experience for our models. Feedback is welcome!

docs/transformers_deploy_guide_cn.md ADDED Viewed

	@@ -0,0 +1,94 @@

+# MiniMax M2.7 模型 Transformers 部署指南
+[英文版](./transformers_deploy_guide.md) | [中文版](./transformers_deploy_guide_cn.md)
+## 本文档适用模型
+本文档适用以下模型，只需在部署时修改模型名称即可。
+- [MiniMaxAI/MiniMax-M2.7](https://huggingface.co/MiniMaxAI/MiniMax-M2.7)
+- [MiniMaxAI/MiniMax-M2.5](https://huggingface.co/MiniMaxAI/MiniMax-M2.5)
+- [MiniMaxAI/MiniMax-M2.1](https://huggingface.co/MiniMaxAI/MiniMax-M2.1)
+- [MiniMaxAI/MiniMax-M2](https://huggingface.co/MiniMaxAI/MiniMax-M2)
+以下以 MiniMax-M2.7 为例说明部署流程。
+## 环境要求
+- OS：Linux
+- Python：3.9 - 3.12
+- Transformers: 4.57.1
+- GPU：
+  - compute capability 7.0 or higher
+  - 显存需求：权重需要 220 GB
+## 使用 Python 部署
+建议使用虚拟环境（如 **venv**、**conda**、**uv**）以避免依赖冲突。
+建议在全新的 Python 环境中安装 Transformers:
+```bash
+uv pip install transformers==4.57.1 torch accelerate --torch-backend=auto
+```
+运行如下 Python 命令运行模型，Transformers 会自动从 Huggingface 下载并缓存 MiniMax-M2.7 模型。
+```python
+from transformers import AutoModelForCausalLM, AutoTokenizer, GenerationConfig
+import torch
+MODEL_PATH = "MiniMaxAI/MiniMax-M2.7"
+model = AutoModelForCausalLM.from_pretrained(
+    MODEL_PATH,
+    device_map="auto",
+    trust_remote_code=True,
+)
+tokenizer = AutoTokenizer.from_pretrained(MODEL_PATH)
+messages = [
+    {"role": "user", "content": [{"type": "text", "text": "What is your favourite condiment?"}]},
+    {"role": "assistant", "content": [{"type": "text", "text": "Well, I'm quite partial to a good squeeze of fresh lemon juice. It adds just the right amount of zesty flavour to whatever I'm cooking up in the kitchen!"}]},
+    {"role": "user", "content": [{"type": "text", "text": "Do you have mayonnaise recipes?"}]}
+]
+model_inputs = tokenizer.apply_chat_template(messages, return_tensors="pt", add_generation_prompt=True).to("cuda")
+generated_ids = model.generate(model_inputs, max_new_tokens=100, generation_config=model.generation_config)
+response = tokenizer.batch_decode(generated_ids)[0]
+print(response)
+```
+## 常见问题
+### Huggingface 网络问题
+如果遇到网络问题，可以设置代理后再进行拉取。
+```bash
+export HF_ENDPOINT=https://hf-mirror.com
+```
+### MiniMax-M2 model is not currently supported
+请确认开启 trust_remote_code=True。
+## 获取支持
+如果在部署 MiniMax 模型过程中遇到任何问题：
+- 通过邮箱 [model@minimax.io](mailto:model@minimax.io) 等官方渠道联系我们的技术支持团队
+- 在我们的 [GitHub](https://github.com/MiniMax-AI) 仓库提交 Issue
+- 通过我们的 [官方企业微信交流群](https://github.com/MiniMax-AI/MiniMax-AI.github.io/blob/main/images/wechat-qrcode.jpeg) 反馈
+我们会持续优化模型的部署体验，欢迎反馈！

docs/vllm_deploy_guide.md ADDED Viewed

	@@ -0,0 +1,118 @@

+# MiniMax M2.7 Model vLLM Deployment Guide
+[English Version](./vllm_deploy_guide.md) | [Chinese Version](./vllm_deploy_guide_cn.md)
+We recommend using [vLLM](https://docs.vllm.ai/en/stable/) to deploy the [MiniMax-M2.7](https://huggingface.co/MiniMaxAI/MiniMax-M2.7) model. vLLM is a high-performance inference engine with excellent serving throughput, efficient and intelligent memory management, powerful batch request processing capabilities, and deeply optimized underlying performance. We recommend reviewing vLLM's official documentation to check hardware compatibility before deployment.
+## Applicable Models
+This document applies to the following models. You only need to change the model name during deployment.
+- [MiniMaxAI/MiniMax-M2.7](https://huggingface.co/MiniMaxAI/MiniMax-M2.7)
+- [MiniMaxAI/MiniMax-M2.5](https://huggingface.co/MiniMaxAI/MiniMax-M2.5)
+- [MiniMaxAI/MiniMax-M2.1](https://huggingface.co/MiniMaxAI/MiniMax-M2.1)
+- [MiniMaxAI/MiniMax-M2](https://huggingface.co/MiniMaxAI/MiniMax-M2)
+The deployment process is illustrated below using MiniMax-M2.7 as an example.
+## System Requirements
+- OS: Linux
+- Python: 3.9 - 3.12
+- GPU:
+  - compute capability 7.0 or higher
+  - Memory requirements: 220 GB for weights, 240 GB per 1M context tokens
+The following are recommended configurations; actual requirements should be adjusted based on your use case:
+- **96G x4** GPU: Supports a total KV Cache capacity of 400K tokens.
+- **144G x8** GPU: Supports a total KV Cache capacity of up to 3M tokens.
+> **Note**: The values above represent the total aggregate hardware KV Cache capacity. The maximum context length per individual sequence remains **196K** tokens.
+## Deployment with Python
+It is recommended to use a virtual environment (such as **venv**, **conda**, or **uv**) to avoid dependency conflicts.
+We recommend installing vLLM in a fresh Python environment:
+```bash
+uv venv
+source .venv/bin/activate
+uv pip install vllm --torch-backend=auto
+```
+Run the following command to start the vLLM server. vLLM will automatically download and cache the MiniMax-M2.7 model from Hugging Face.
+4-GPU deployment command:
+```bash
+SAFETENSORS_FAST_GPU=1 vllm serve \
+    MiniMaxAI/MiniMax-M2.7 --trust-remote-code \
+    --tensor-parallel-size 4 \
+    --enable-auto-tool-choice --tool-call-parser minimax_m2 \
+    --reasoning-parser minimax_m2_append_think
+```
+8-GPU deployment command:
+```bash
+SAFETENSORS_FAST_GPU=1 vllm serve \
+    MiniMaxAI/MiniMax-M2.7 --trust-remote-code \
+    --enable_expert_parallel --tensor-parallel-size 8 \
+    --enable-auto-tool-choice --tool-call-parser minimax_m2 \
+    --reasoning-parser minimax_m2_append_think
+```
+## Testing Deployment
+After startup, you can test the vLLM OpenAI-compatible API with the following command:
+```bash
+curl http://localhost:8000/v1/chat/completions \
+    -H "Content-Type: application/json" \
+    -d '{
+        "model": "MiniMaxAI/MiniMax-M2.7",
+        "messages": [
+            {"role": "system", "content": [{"type": "text", "text": "You are a helpful assistant."}]},
+            {"role": "user", "content": [{"type": "text", "text": "Who won the world series in 2020?"}]}
+        ]
+    }'
+```
+## Common Issues
+### MiniMax-M2 model is not currently supported
+This vLLM version is outdated. Please upgrade to the latest version.
+### torch.AcceleratorError: CUDA error: an illegal memory access was encountered
+Add `--compilation-config "{\"cudagraph_mode\": \"PIECEWISE\"}"` to the startup parameters to resolve this issue. For example:
+```bash
+SAFETENSORS_FAST_GPU=1 vllm serve \
+    MiniMaxAI/MiniMax-M2.7 --trust-remote-code \
+    --enable_expert_parallel --tensor-parallel-size 8 \
+    --enable-auto-tool-choice --tool-call-parser minimax_m2 \
+    --reasoning-parser minimax_m2_append_think \
+    --compilation-config "{\"cudagraph_mode\": \"PIECEWISE\"}"
+```
+### Output is garbled
+If you encounter corrupted output when using vLLM to serve these models, you can upgrade to the nightly version (ensure it is a version after commit [cf3eacfe58fa9e745c2854782ada884a9f992cf7](https://github.com/vllm-project/vllm/commit/cf3eacfe58fa9e745c2854782ada884a9f992cf7))
+## Getting Support
+If you encounter any issues while deploying the MiniMax model:
+- Contact our technical support team through official channels such as email at [model@minimax.io](mailto:model@minimax.io)
+- Submit an issue on our [GitHub](https://github.com/MiniMax-AI) repository
+We continuously optimize the deployment experience for our models. Feedback is welcome!

docs/vllm_deploy_guide_cn.md ADDED Viewed

	@@ -0,0 +1,128 @@

+# MiniMax M2.7 模型 vLLM 部署指南
+[英文版](./vllm_deploy_guide.md) | [中文版](./vllm_deploy_guide_cn.md)
+我们推荐使用 [vLLM](https://docs.vllm.ai/en/stable/) 来部署 [MiniMax-M2.7](https://huggingface.co/MiniMaxAI/MiniMax-M2.7) 模型。vLLM 是一个高性能的推理引擎，其具有卓越的服务吞吐、高效智能的内存管理机制、强大的批量请求处理能力、深度优化的底层性能等特性。我们建议在部署之前查看 vLLM 的官方文档以检查硬件兼容性。
+## 本文档适用模型
+本文档适用以下模型，只需在部署时修改模型名称即可。
+- [MiniMaxAI/MiniMax-M2.7](https://huggingface.co/MiniMaxAI/MiniMax-M2.7)
+- [MiniMaxAI/MiniMax-M2.5](https://huggingface.co/MiniMaxAI/MiniMax-M2.5)
+- [MiniMaxAI/MiniMax-M2.1](https://huggingface.co/MiniMaxAI/MiniMax-M2.1)
+- [MiniMaxAI/MiniMax-M2](https://huggingface.co/MiniMaxAI/MiniMax-M2)
+以下以 MiniMax-M2.7 为例说明部署流程。
+## 环境要求
+- OS：Linux
+- Python：3.9 - 3.12
+- GPU：
+  - compute capability 7.0 or higher
+  - 显存需求：权重需要 220 GB，每 1M 上下文 token 需要 240 GB
+以下为推荐配置，实际需求请根据业务场景调整：
+- **96G x4 GPU**：总 KV Cache 容量支持 40 万 token。
+- **144G x8 GPU**：总 KV Cache 容量支持高达 300 万 token。
+> **注**：以上数值为硬件支持的最大并发缓存总量，模型单序列（Single Sequence）长度上限仍为 196k。
+## 使用 Python 部署
+建议使用虚拟环境（如 **venv**、**conda**、**uv**）以避免依赖冲突。
+建议在全新的 Python 环境中安装 vLLM：
+```bash
+uv venv
+source .venv/bin/activate
+uv pip install vllm --torch-backend=auto
+```
+运行如下命令启动 vLLM 服务器，vLLM 会自动从 Huggingface 下载并缓存 MiniMax-M2.7 模型。
+4 卡部署命令：
+```bash
+SAFETENSORS_FAST_GPU=1 vllm serve \
+    MiniMaxAI/MiniMax-M2.7 --trust-remote-code \
+    --tensor-parallel-size 4 \
+    --enable-auto-tool-choice --tool-call-parser minimax_m2 \
+    --reasoning-parser minimax_m2_append_think
+```
+8 卡部署命令：
+```bash
+SAFETENSORS_FAST_GPU=1 vllm serve \
+    MiniMaxAI/MiniMax-M2.7 --trust-remote-code \
+    --enable_expert_parallel --tensor-parallel-size 8 \
+    --enable-auto-tool-choice --tool-call-parser minimax_m2 \
+    --reasoning-parser minimax_m2_append_think
+```
+## 测试部署
+启动后，可以通过如下命令测试 vLLM OpenAI 兼容接口：
+```bash
+curl http://localhost:8000/v1/chat/completions \
+    -H "Content-Type: application/json" \
+    -d '{
+        "model": "MiniMaxAI/MiniMax-M2.7",
+        "messages": [
+            {"role": "system", "content": [{"type": "text", "text": "You are a helpful assistant."}]},
+            {"role": "user", "content": [{"type": "text", "text": "Who won the world series in 2020?"}]}
+        ]
+    }'
+```
+## 常见问题
+### Huggingface 网络问题
+如果遇到网络问题，可以设置代理后再进行拉取。
+```bash
+export HF_ENDPOINT=https://hf-mirror.com
+```
+### MiniMax-M2 model is not currently supported
+该 vLLM 版本过旧，请升级到最新版本。
+### torch.AcceleratorError: CUDA error: an illegal memory access was encountered
+在启动参数添加 `--compilation-config "{\"cudagraph_mode\": \"PIECEWISE\"}"` 可以解决。例如：
+```bash
+SAFETENSORS_FAST_GPU=1 vllm serve \
+    MiniMaxAI/MiniMax-M2.7 --trust-remote-code \
+    --enable_expert_parallel --tensor-parallel-size 8 \
+    --enable-auto-tool-choice --tool-call-parser minimax_m2 \
+    --reasoning-parser minimax_m2_append_think \
+    --compilation-config "{\"cudagraph_mode\": \"PIECEWISE\"}"
+```
+### 模型输出乱码
+如果您在使用 vLLM 运行这些模型时遇到输出乱码，可以升级到最新版本（请至少确保版本在提交 [cf3eacfe58fa9e745c2854782ada884a9f992cf7](https://github.com/vllm-project/vllm/commit/cf3eacfe58fa9e745c2854782ada884a9f992cf7) 之后）。
+## 获取支持
+如果在部署 MiniMax 模型过程中遇到任何问题：
+- 通过邮箱 [model@minimax.io](mailto:model@minimax.io) 等官方渠道联系我们的技术支持团队
+- 在我们的 [GitHub](https://github.com/MiniMax-AI) 仓库提交 Issue
+- 通过我们的 [官方企业微信交流群](https://github.com/MiniMax-AI/MiniMax-AI.github.io/blob/main/images/wechat-qrcode.jpeg) 反馈
+我们会持续优化模型的部署体验，欢迎反馈！

figures/agent_harness.png ADDED Viewed

Git LFS Details

SHA256: 7c661c39ff84fcc10a000f0dbc3b648e22bda4c8ebb43ce30cee1bc3d21c6c01
Pointer size: 131 Bytes
Size of remote file: 312 kB

figures/agent_teams.gif ADDED Viewed

Git LFS Details

SHA256: 2f72c28868b6f5cc641435923a7340e4b7a8e9ecf79cd46ca31b277152a0dfd1
Pointer size: 132 Bytes
Size of remote file: 7.44 MB

figures/banner.png ADDED Viewed

Git LFS Details

SHA256: d524f03ea8db52076ed29a070526c170ecf5f20db789d4e79cf92521234758d8
Pointer size: 131 Bytes
Size of remote file: 120 kB

figures/benchmark_overview.png ADDED Viewed

figures/mle_bench.png ADDED Viewed

Git LFS Details

SHA256: 6bdbff9fc7f90735f75bd51ecd1eb1a1db8b4804392744c4e8544988bdc7978c
Pointer size: 131 Bytes
Size of remote file: 123 kB

generation_config.json ADDED Viewed

	@@ -0,0 +1,9 @@

+{
+  "bos_token_id": 200019,
+  "do_sample": true,
+  "eos_token_id": 200020,
+  "temperature": 1.0,
+  "top_p": 0.95,
+  "top_k": 40,
+  "transformers_version": "4.46.1"
+}

merges.txt ADDED Viewed

The diff for this file is too large to render. See raw diff

model-00000-of-00130.safetensors ADDED Viewed

	@@ -0,0 +1,3 @@

+version https://git-lfs.github.com/spec/v1
+oid sha256:9785f5a87c85710e38f4ca11f819f3d137ff84615af1bc0ba533b94681addf27
+size 3693062744

model-00001-of-00130.safetensors ADDED Viewed

	@@ -0,0 +1,3 @@

+version https://git-lfs.github.com/spec/v1
+oid sha256:d2ed94efe077a4498b788706e059d82780deb54436a70a5a9664b716d6cdc83e
+size 1208321176

model-00002-of-00130.safetensors ADDED Viewed

	@@ -0,0 +1,3 @@

+version https://git-lfs.github.com/spec/v1
+oid sha256:f0c1b97aff37136b5d89a9df22acf7109fa824ccef5f9ff4f763b7869dfc5650
+size 2463868936

model-00003-of-00130.safetensors ADDED Viewed

	@@ -0,0 +1,3 @@

+version https://git-lfs.github.com/spec/v1
+oid sha256:93be479ff1b6912ff1a7e54f4c4a4e4d67124d1811df8e39d50b981b1b43d8e6
+size 1208321176

model-00004-of-00130.safetensors ADDED Viewed

	@@ -0,0 +1,3 @@

+version https://git-lfs.github.com/spec/v1
+oid sha256:5d5bead700b8f82dd2a50cee205c37f5642020c414452869693da06df384a9eb
+size 2463868936

model-00005-of-00130.safetensors ADDED Viewed

	@@ -0,0 +1,3 @@

+version https://git-lfs.github.com/spec/v1
+oid sha256:99444d6d83c614776397faa167dc908d48016414e0dd6edef57fd9c040e01d21
+size 1208321176

model-00006-of-00130.safetensors ADDED Viewed

	@@ -0,0 +1,3 @@

+version https://git-lfs.github.com/spec/v1
+oid sha256:df42d1d91b84ed41f846775a274dbd382185fdf7595009dcd016bd805e25eb1b
+size 2463868936

model-00007-of-00130.safetensors ADDED Viewed

	@@ -0,0 +1,3 @@

+version https://git-lfs.github.com/spec/v1
+oid sha256:18882ffcb4f2dddfe6b8766393c68208b524aa4520ed921234a66b11548440eb
+size 1208321176

model-00008-of-00130.safetensors ADDED Viewed

	@@ -0,0 +1,3 @@

+version https://git-lfs.github.com/spec/v1
+oid sha256:cf8ead5d7b01543a3fafc5a39240b1a3d9fe1cf25b360eb99e7a751359db9705
+size 2463868936

model-00009-of-00130.safetensors ADDED Viewed

	@@ -0,0 +1,3 @@

+version https://git-lfs.github.com/spec/v1
+oid sha256:d897820ce912aa7ae2feb4377d9b8684eca38c18be550b6bcf7316cb9d7c6e30
+size 1208321176

model-00010-of-00130.safetensors ADDED Viewed

	@@ -0,0 +1,3 @@

+version https://git-lfs.github.com/spec/v1
+oid sha256:734eee6e62863c518a976d41b6c4122ed974cf87e52cd2d7e7df0187a3141b87
+size 2463868936

model-00011-of-00130.safetensors ADDED Viewed

	@@ -0,0 +1,3 @@

+version https://git-lfs.github.com/spec/v1
+oid sha256:1237cbe1b9915bfda1efb8ced7d5a4266a0083a3b4c3fa401c4a003e3fea20fd
+size 1208321176

model-00012-of-00130.safetensors ADDED Viewed

	@@ -0,0 +1,3 @@

+version https://git-lfs.github.com/spec/v1
+oid sha256:069b272af35289d3c499e98f867b1ffecb1f96980c583bf77b1d4d23c8b7a713
+size 2463868936

model-00013-of-00130.safetensors ADDED Viewed

	@@ -0,0 +1,3 @@

+version https://git-lfs.github.com/spec/v1
+oid sha256:045403b45c8951c3ea3c68b288f04255e0e2fc4de47293f9b941964212b8253e
+size 1208321176

model-00014-of-00130.safetensors ADDED Viewed

	@@ -0,0 +1,3 @@

+version https://git-lfs.github.com/spec/v1
+oid sha256:0277da3d1063a00618b32992617a2448c95c850c1f26dc4024d70ae920a35a25
+size 2463868936

model-00015-of-00130.safetensors ADDED Viewed

	@@ -0,0 +1,3 @@

+version https://git-lfs.github.com/spec/v1
+oid sha256:d2a9db97dbab9f2a324219d4ba019656b6b635fae3b868d7f2a4fd6e3bab5e66
+size 1208321176

model-00016-of-00130.safetensors ADDED Viewed

	@@ -0,0 +1,3 @@

+version https://git-lfs.github.com/spec/v1
+oid sha256:90776eaf143864ecb632c059fefd4167e27c5644ba4eb50d65afa5291cff666e
+size 2463868936

model-00017-of-00130.safetensors ADDED Viewed

	@@ -0,0 +1,3 @@

+version https://git-lfs.github.com/spec/v1
+oid sha256:4ea50b70dae5f8b55b1990a6b6cad9291349b45162548e9d48d63b2a144e3c23
+size 1208321176

model-00018-of-00130.safetensors ADDED Viewed

	@@ -0,0 +1,3 @@

+version https://git-lfs.github.com/spec/v1
+oid sha256:2a239e9eae27174937d5547d8e5e743e84bd7eaea50390510e4cd8f15511447b
+size 2463868936

model-00019-of-00130.safetensors ADDED Viewed

	@@ -0,0 +1,3 @@

+version https://git-lfs.github.com/spec/v1
+oid sha256:5e041358d2ce0d92517b13508046baf08807d46adb33dda5d23728a4cef45f2b
+size 1208321176

model-00020-of-00130.safetensors ADDED Viewed

	@@ -0,0 +1,3 @@

+version https://git-lfs.github.com/spec/v1
+oid sha256:4f4f7af9ded3e7d5775012eae2c7dee63518c799ebbe42a47949aa7f560c5f43
+size 2463869968

model-00021-of-00130.safetensors ADDED Viewed

	@@ -0,0 +1,3 @@

+version https://git-lfs.github.com/spec/v1
+oid sha256:8a76ddac05820e58676b3b56e2990c598dae551f1f65adf55a90a3754f66e2b4
+size 1208321688

model-00022-of-00130.safetensors ADDED Viewed

	@@ -0,0 +1,3 @@

+version https://git-lfs.github.com/spec/v1
+oid sha256:c080ad8c3b5032434973e205a074e4d1a41edd399a383dc1c6d80ebb073ca09e
+size 2463869968

model-00023-of-00130.safetensors ADDED Viewed

	@@ -0,0 +1,3 @@

+version https://git-lfs.github.com/spec/v1
+oid sha256:9eee017222d3eb90afa5126fccb194de12c67828bd4353b3a466ce3da17877d2
+size 1208321688

model-00024-of-00130.safetensors ADDED Viewed

	@@ -0,0 +1,3 @@

+version https://git-lfs.github.com/spec/v1
+oid sha256:e3d3c543000e2fd6180bb17c289f36e46256bf0c76f7ae98a7087eb4264db605
+size 2463869968

model-00025-of-00130.safetensors ADDED Viewed

	@@ -0,0 +1,3 @@

+version https://git-lfs.github.com/spec/v1
+oid sha256:68580bdb4da65c22fb95a16e7fe13b1f0bbde861327d7c0bb6cb76a86794d38d
+size 1208321688

model-00026-of-00130.safetensors ADDED Viewed

	@@ -0,0 +1,3 @@

+version https://git-lfs.github.com/spec/v1
+oid sha256:c0ca69318b53d7ec6f7fcfa7981ed2ec402e73302fd5ea62ed77311f4eb8be73
+size 2463869968

model-00027-of-00130.safetensors ADDED Viewed

	@@ -0,0 +1,3 @@

+version https://git-lfs.github.com/spec/v1
+oid sha256:a6f03ff04b01299dceaf26fe0a0a503d6e0abc58eba94e8796e933e40bd10a5e
+size 1208321688

model-00028-of-00130.safetensors ADDED Viewed

	@@ -0,0 +1,3 @@

+version https://git-lfs.github.com/spec/v1
+oid sha256:6432450282a2cd79475b57bf5b83380addf0b8d36586c750bc4fbf37ce04af6e
+size 2463869968

model-00029-of-00130.safetensors ADDED Viewed

	@@ -0,0 +1,3 @@

+version https://git-lfs.github.com/spec/v1
+oid sha256:961ca8675f7ee7a1a65e5ea5f1e35dfe7427d566e68a1f56f04a463252763683
+size 1208321688