Skip to content

Commit 15b6e06

Browse files
qinxuyexiaoyesoso
authored andcommitted
chore(docs): update generated model docs
1 parent e6853b3 commit 15b6e06

10 files changed

Lines changed: 102 additions & 117 deletions

File tree

doc/source/getting_started/installation.rst

Lines changed: 2 additions & 1 deletion
Original file line numberDiff line numberDiff line change
@@ -65,7 +65,7 @@ Currently, supported models include:
6565

6666
.. vllm_start
6767
68-
- ``code-llama``, ``code-llama-instruct``, ``code-llama-python``, ``deepseek``, ``deepseek-chat``, ``deepseek-coder``, ``deepseek-coder-instruct``, ``deepseek-r1-distill-llama``, ``gorilla-openfunctions-v2``, ``HuatuoGPT-o1-LLaMA-3.1``, ``llama-2``, ``llama-2-chat``, ``llama-3``, ``llama-3-instruct``, ``llama-3.1``, ``llama-3.1-instruct``, ``llama-3.3-instruct``, ``tiny-llama``, ``wizardcoder-python-v1.0``, ``wizardmath-v1.0``, ``Yi``, ``Yi-1.5``, ``Yi-1.5-chat``, ``Yi-1.5-chat-16k``, ``Yi-200k``, ``Yi-chat``
68+
- ``code-llama``, ``code-llama-instruct``, ``code-llama-python``, ``deepseek``, ``deepseek-chat``, ``deepseek-coder``, ``deepseek-coder-instruct``, ``deepseek-r1-distill-llama``, ``gorilla-openfunctions-v2``, ``HuatuoGPT-o1-LLaMA-3.1``, ``llama-2``, ``llama-2-chat``, ``llama-3``, ``llama-3-instruct``, ``llama-3.1``, ``llama-3.1-instruct``, ``llama-3.3-instruct``, ``minicpm5-1b``, ``tiny-llama``, ``wizardcoder-python-v1.0``, ``wizardmath-v1.0``, ``Yi``, ``Yi-1.5``, ``Yi-1.5-chat``, ``Yi-1.5-chat-16k``, ``Yi-200k``, ``Yi-chat``
6969
- ``codestral-v0.1``, ``mistral-instruct-v0.1``, ``mistral-instruct-v0.2``, ``mistral-instruct-v0.3``, ``mistral-large-instruct``, ``mistral-nemo-instruct``, ``mistral-v0.1``, ``openhermes-2.5``, ``seallm_v2``
7070
- ``Baichuan-M2``, ``codeqwen1.5``, ``codeqwen1.5-chat``, ``deepseek-r1-distill-qwen``, ``DianJin-R1``, ``fin-r1``, ``HuatuoGPT-o1-Qwen2.5``, ``KAT-V1``, ``marco-o1``, ``qwen1.5-chat``, ``qwen2-instruct``, ``qwen2.5``, ``qwen2.5-coder``, ``qwen2.5-coder-instruct``, ``qwen2.5-instruct``, ``qwen2.5-instruct-1m``, ``qwenLong-l1``, ``QwQ-32B``, ``QwQ-32B-Preview``, ``seallms-v3``, ``skywork-or1``, ``skywork-or1-preview``, ``XiYanSQL-QwenCoder-2504``
7171
- ``llama-3.2-vision``, ``llama-3.2-vision-instruct``
@@ -98,6 +98,7 @@ Currently, supported models include:
9898
- ``MiniMax-M2``, ``MiniMax-M2.5``, ``MiniMax-M2.7``
9999
- ``GLM-4.7-Flash``
100100
- ``glm-5``, ``glm-5.1``
101+
- ``DeepSeek-V4-Flash``, ``DeepSeek-V4-Pro``
101102
.. vllm_end
102103
103104
To install Xinference and vLLM::

doc/source/models/builtin/llm/deepseek-v4-flash.rst

Lines changed: 1 addition & 1 deletion
Original file line numberDiff line numberDiff line change
@@ -20,7 +20,7 @@ Model Spec 1 (pytorch, 284 Billion)
2020
- **Model Format:** pytorch
2121
- **Model Size (in billions):** 284
2222
- **Quantizations:** none
23-
- **Engines**: Transformers
23+
- **Engines**: vLLM, Transformers
2424
- **Model ID:** deepseek-ai/DeepSeek-V4-Flash
2525
- **Model Hubs**: `Hugging Face <https://huggingface.co/deepseek-ai/DeepSeek-V4-Flash>`__, `ModelScope <https://modelscope.cn/models/deepseek-ai/DeepSeek-V4-Flash>`__
2626

doc/source/models/builtin/llm/deepseek-v4-pro.rst

Lines changed: 1 addition & 1 deletion
Original file line numberDiff line numberDiff line change
@@ -20,7 +20,7 @@ Model Spec 1 (pytorch, 1600 Billion)
2020
- **Model Format:** pytorch
2121
- **Model Size (in billions):** 1600
2222
- **Quantizations:** none
23-
- **Engines**: Transformers
23+
- **Engines**: vLLM, Transformers
2424
- **Model ID:** deepseek-ai/DeepSeek-V4-Pro
2525
- **Model Hubs**: `Hugging Face <https://huggingface.co/deepseek-ai/DeepSeek-V4-Pro>`__, `ModelScope <https://modelscope.cn/models/deepseek-ai/DeepSeek-V4-Pro>`__
2626

doc/source/models/builtin/llm/dianjin-r1.rst

Lines changed: 1 addition & 1 deletion
Original file line numberDiff line numberDiff line change
@@ -7,7 +7,7 @@ DianJin-R1
77
- **Context Length:** 32768
88
- **Model Name:** DianJin-R1
99
- **Languages:** en, zh
10-
- **Abilities:** chat, tools
10+
- **Abilities:** chat, reasoning, hybrid, tools
1111
- **Description:** Tongyi DianJin is a financial intelligence solution platform built by Alibaba Cloud, dedicated to providing financial business developers with a convenient artificial intelligence application development environment.
1212

1313
Specifications

doc/source/models/builtin/llm/index.rst

Lines changed: 9 additions & 3 deletions
Original file line numberDiff line numberDiff line change
@@ -187,7 +187,7 @@ The following is a list of built-in LLM in Xinference:
187187
- DeepSeek-VL2, an advanced series of large Mixture-of-Experts (MoE) Vision-Language Models that significantly improves upon its predecessor, DeepSeek-VL. DeepSeek-VL2 demonstrates superior capabilities across various tasks, including but not limited to visual question answering, optical character recognition, document/table/chart understanding, and visual grounding.
188188

189189
* - :ref:`dianjin-r1 <models_llm_dianjin-r1>`
190-
- chat, tools
190+
- chat, reasoning, hybrid, tools
191191
- 32768
192192
- Tongyi DianJin is a financial intelligence solution platform built by Alibaba Cloud, dedicated to providing financial business developers with a convenient artificial intelligence application development environment.
193193

@@ -426,6 +426,11 @@ The following is a list of built-in LLM in Xinference:
426426
- 32768
427427
- MiniCPM4 series are highly efficient large language models (LLMs) designed explicitly for end-side devices, which achieves this efficiency through systematic innovation in four key dimensions: model architecture, training data, training algorithms, and inference systems.
428428

429+
* - :ref:`minicpm5-1b <models_llm_minicpm5-1b>`
430+
- chat, reasoning, hybrid, tools
431+
- 131072
432+
- MiniCPM5-1B is the first model in the MiniCPM5 series. It is a dense 1B Transformer built for on-device, local deployment, and resource-constrained scenarios, reaching 1B-class open-source SOTA. Supports hybrid thinking via enable_thinking and native XML-style tool calling (MCP-compatible).
433+
429434
* - :ref:`minimax-m2 <models_llm_minimax-m2>`
430435
- chat, tools, reasoning
431436
- 196608
@@ -785,7 +790,7 @@ The following is a list of built-in LLM in Xinference:
785790
.. toctree::
786791
:maxdepth: 3
787792

788-
793+
789794
baichuan-2
790795

791796
baichuan-2-chat
@@ -950,6 +955,8 @@ The following is a list of built-in LLM in Xinference:
950955

951956
minicpm4
952957

958+
minicpm5-1b
959+
953960
minimax-m2
954961

955962
minimax-m2.5
@@ -1092,4 +1099,3 @@ The following is a list of built-in LLM in Xinference:
10921099

10931100
yi-chat
10941101

1095-
Lines changed: 62 additions & 0 deletions
Original file line numberDiff line numberDiff line change
@@ -0,0 +1,62 @@
1+
.. _models_llm_minicpm5-1b:
2+
3+
========================================
4+
minicpm5-1b
5+
========================================
6+
7+
- **Context Length:** 131072
8+
- **Model Name:** minicpm5-1b
9+
- **Languages:** en, zh
10+
- **Abilities:** chat, reasoning, hybrid, tools
11+
- **Description:** MiniCPM5-1B is the first model in the MiniCPM5 series. It is a dense 1B Transformer built for on-device, local deployment, and resource-constrained scenarios, reaching 1B-class open-source SOTA. Supports hybrid thinking via enable_thinking and native XML-style tool calling (MCP-compatible).
12+
13+
Specifications
14+
^^^^^^^^^^^^^^
15+
16+
17+
Model Spec 1 (pytorch, 1 Billion)
18+
++++++++++++++++++++++++++++++++++++++++
19+
20+
- **Model Format:** pytorch
21+
- **Model Size (in billions):** 1
22+
- **Quantizations:** none
23+
- **Engines**: vLLM, Transformers, SGLang
24+
- **Model ID:** openbmb/MiniCPM5-1B
25+
- **Model Hubs**: `Hugging Face <https://huggingface.co/openbmb/MiniCPM5-1B>`__, `ModelScope <https://modelscope.cn/models/OpenBMB/MiniCPM5-1B>`__
26+
27+
Execute the following command to launch the model, remember to replace ``${quantization}`` with your
28+
chosen quantization method from the options listed above::
29+
30+
xinference launch --model-engine ${engine} --model-name minicpm5-1b --size-in-billions 1 --model-format pytorch --quantization ${quantization}
31+
32+
33+
Model Spec 2 (ggufv2, 1 Billion)
34+
++++++++++++++++++++++++++++++++++++++++
35+
36+
- **Model Format:** ggufv2
37+
- **Model Size (in billions):** 1
38+
- **Quantizations:** F16, Q4_K_M, Q8_0
39+
- **Engines**: vLLM, llama.cpp
40+
- **Model ID:** openbmb/MiniCPM5-1B-GGUF
41+
- **Model Hubs**: `Hugging Face <https://huggingface.co/openbmb/MiniCPM5-1B-GGUF>`__, `ModelScope <https://modelscope.cn/models/OpenBMB/MiniCPM5-1B-GGUF>`__
42+
43+
Execute the following command to launch the model, remember to replace ``${quantization}`` with your
44+
chosen quantization method from the options listed above::
45+
46+
xinference launch --model-engine ${engine} --model-name minicpm5-1b --size-in-billions 1 --model-format ggufv2 --quantization ${quantization}
47+
48+
49+
Model Spec 3 (mlx, 1 Billion)
50+
++++++++++++++++++++++++++++++++++++++++
51+
52+
- **Model Format:** mlx
53+
- **Model Size (in billions):** 1
54+
- **Quantizations:** 4bit
55+
- **Engines**: MLX
56+
- **Model ID:** openbmb/MiniCPM5-1B-MLX
57+
- **Model Hubs**: `Hugging Face <https://huggingface.co/openbmb/MiniCPM5-1B-MLX>`__, `ModelScope <https://modelscope.cn/models/OpenBMB/MiniCPM5-1B-MLX>`__
58+
59+
Execute the following command to launch the model, remember to replace ``${quantization}`` with your
60+
chosen quantization method from the options listed above::
61+
62+
xinference launch --model-engine ${engine} --model-name minicpm5-1b --size-in-billions 1 --model-format mlx --quantization ${quantization}

doc/source/models/builtin/llm/qwen3.6.rst

Lines changed: 21 additions & 5 deletions
Original file line numberDiff line numberDiff line change
@@ -46,7 +46,23 @@ chosen quantization method from the options listed above::
4646
xinference launch --model-engine ${engine} --model-name qwen3.6 --size-in-billions 27 --model-format fp8 --quantization ${quantization}
4747

4848

49-
Model Spec 3 (pytorch, 35 Billion)
49+
Model Spec 3 (awq, 27 Billion)
50+
++++++++++++++++++++++++++++++++++++++++
51+
52+
- **Model Format:** awq
53+
- **Model Size (in billions):** 27
54+
- **Quantizations:** Int4
55+
- **Engines**: vLLM, Transformers, SGLang
56+
- **Model ID:** QuantTrio/Qwen3.6-27B-AWQ
57+
- **Model Hubs**: `Hugging Face <https://huggingface.co/QuantTrio/Qwen3.6-27B-AWQ>`__, `ModelScope <https://modelscope.cn/models/tclf90/Qwen3.6-27B-AWQ>`__
58+
59+
Execute the following command to launch the model, remember to replace ``${quantization}`` with your
60+
chosen quantization method from the options listed above::
61+
62+
xinference launch --model-engine ${engine} --model-name qwen3.6 --size-in-billions 27 --model-format awq --quantization ${quantization}
63+
64+
65+
Model Spec 4 (pytorch, 35 Billion)
5066
++++++++++++++++++++++++++++++++++++++++
5167

5268
- **Model Format:** pytorch
@@ -62,7 +78,7 @@ chosen quantization method from the options listed above::
6278
xinference launch --model-engine ${engine} --model-name qwen3.6 --size-in-billions 35 --model-format pytorch --quantization ${quantization}
6379

6480

65-
Model Spec 4 (fp8, 35 Billion)
81+
Model Spec 5 (fp8, 35 Billion)
6682
++++++++++++++++++++++++++++++++++++++++
6783

6884
- **Model Format:** fp8
@@ -78,7 +94,7 @@ chosen quantization method from the options listed above::
7894
xinference launch --model-engine ${engine} --model-name qwen3.6 --size-in-billions 35 --model-format fp8 --quantization ${quantization}
7995

8096

81-
Model Spec 5 (awq, 35 Billion)
97+
Model Spec 6 (awq, 35 Billion)
8298
++++++++++++++++++++++++++++++++++++++++
8399

84100
- **Model Format:** awq
@@ -94,7 +110,7 @@ chosen quantization method from the options listed above::
94110
xinference launch --model-engine ${engine} --model-name qwen3.6 --size-in-billions 35 --model-format awq --quantization ${quantization}
95111

96112

97-
Model Spec 6 (ggufv2, 35 Billion)
113+
Model Spec 7 (ggufv2, 35 Billion)
98114
++++++++++++++++++++++++++++++++++++++++
99115

100116
- **Model Format:** ggufv2
@@ -110,7 +126,7 @@ chosen quantization method from the options listed above::
110126
xinference launch --model-engine ${engine} --model-name qwen3.6 --size-in-billions 35 --model-format ggufv2 --quantization ${quantization}
111127

112128

113-
Model Spec 7 (mlx, 35 Billion)
129+
Model Spec 8 (mlx, 35 Billion)
114130
++++++++++++++++++++++++++++++++++++++++
115131

116132
- **Model Format:** mlx

doc/source/user_guide/backends.rst

Lines changed: 2 additions & 1 deletion
Original file line numberDiff line numberDiff line change
@@ -128,7 +128,7 @@ Currently, supported model includes:
128128

129129
.. vllm_start
130130
131-
- ``code-llama``, ``code-llama-instruct``, ``code-llama-python``, ``deepseek``, ``deepseek-chat``, ``deepseek-coder``, ``deepseek-coder-instruct``, ``deepseek-r1-distill-llama``, ``gorilla-openfunctions-v2``, ``HuatuoGPT-o1-LLaMA-3.1``, ``llama-2``, ``llama-2-chat``, ``llama-3``, ``llama-3-instruct``, ``llama-3.1``, ``llama-3.1-instruct``, ``llama-3.3-instruct``, ``tiny-llama``, ``wizardcoder-python-v1.0``, ``wizardmath-v1.0``, ``Yi``, ``Yi-1.5``, ``Yi-1.5-chat``, ``Yi-1.5-chat-16k``, ``Yi-200k``, ``Yi-chat``
131+
- ``code-llama``, ``code-llama-instruct``, ``code-llama-python``, ``deepseek``, ``deepseek-chat``, ``deepseek-coder``, ``deepseek-coder-instruct``, ``deepseek-r1-distill-llama``, ``gorilla-openfunctions-v2``, ``HuatuoGPT-o1-LLaMA-3.1``, ``llama-2``, ``llama-2-chat``, ``llama-3``, ``llama-3-instruct``, ``llama-3.1``, ``llama-3.1-instruct``, ``llama-3.3-instruct``, ``minicpm5-1b``, ``tiny-llama``, ``wizardcoder-python-v1.0``, ``wizardmath-v1.0``, ``Yi``, ``Yi-1.5``, ``Yi-1.5-chat``, ``Yi-1.5-chat-16k``, ``Yi-200k``, ``Yi-chat``
132132
- ``codestral-v0.1``, ``mistral-instruct-v0.1``, ``mistral-instruct-v0.2``, ``mistral-instruct-v0.3``, ``mistral-large-instruct``, ``mistral-nemo-instruct``, ``mistral-v0.1``, ``openhermes-2.5``, ``seallm_v2``
133133
- ``Baichuan-M2``, ``codeqwen1.5``, ``codeqwen1.5-chat``, ``deepseek-r1-distill-qwen``, ``DianJin-R1``, ``fin-r1``, ``HuatuoGPT-o1-Qwen2.5``, ``KAT-V1``, ``marco-o1``, ``qwen1.5-chat``, ``qwen2-instruct``, ``qwen2.5``, ``qwen2.5-coder``, ``qwen2.5-coder-instruct``, ``qwen2.5-instruct``, ``qwen2.5-instruct-1m``, ``qwenLong-l1``, ``QwQ-32B``, ``QwQ-32B-Preview``, ``seallms-v3``, ``skywork-or1``, ``skywork-or1-preview``, ``XiYanSQL-QwenCoder-2504``
134134
- ``llama-3.2-vision``, ``llama-3.2-vision-instruct``
@@ -161,6 +161,7 @@ Currently, supported model includes:
161161
- ``MiniMax-M2``, ``MiniMax-M2.5``, ``MiniMax-M2.7``
162162
- ``GLM-4.7-Flash``
163163
- ``glm-5``, ``glm-5.1``
164+
- ``DeepSeek-V4-Flash``, ``DeepSeek-V4-Pro``
164165
.. vllm_end
165166
166167
.. _sglang_backend:

xinference/__init__.py

Lines changed: 0 additions & 82 deletions
Original file line numberDiff line numberDiff line change
@@ -37,88 +37,6 @@ def _install():
3737
default_router = Router.get_instance_or_empty()
3838
Router.set_instance(default_router)
3939

40-
# Compatibility shim: xoscar <= 0.9.6 (the latest PyPI release at the
41-
# time of writing) does not expose ``wait_for``. xinference internals
42-
# (e.g. core/worker.py, core/supervisor.py) call ``xo.wait_for`` for
43-
# actor-aware timeouts that survive cancellation routed through Xoscar
44-
# actor calls -- so we cannot simply alias it to ``asyncio.wait_for``
45-
# (that would re-introduce the timeout-hang bug). Backport the
46-
# actor-aware implementation from xoscar main instead.
47-
import xoscar as _xo
48-
49-
if not hasattr(_xo, "wait_for"):
50-
import asyncio as _asyncio
51-
import logging as _logging
52-
from typing import Any as _Any
53-
from typing import Awaitable as _Awaitable
54-
55-
_wait_for_logger = _logging.getLogger("xinference.compat.xoscar")
56-
57-
async def _xo_wait_for(
58-
fut: "_Awaitable[_Any]", timeout: "float | None" = None
59-
) -> "_Any":
60-
# Mirrors xoscar.api.wait_for: asyncio.wait_for on an actor call
61-
# can hang because the cancellation is delivered through a
62-
# CancelMessage that may never be processed; this wrapper makes
63-
# the timeout fire regardless.
64-
loop = _asyncio.get_running_loop()
65-
new_fut: "_asyncio.Future[_Any]" = loop.create_future()
66-
task = _asyncio.ensure_future(fut)
67-
68-
def _on_done(f: "_asyncio.Future[_Any]") -> None:
69-
if new_fut.done():
70-
return
71-
if f.cancelled():
72-
new_fut.cancel()
73-
elif f.exception() is not None:
74-
new_fut.set_exception(f.exception()) # type: ignore[arg-type]
75-
else:
76-
new_fut.set_result(f.result())
77-
78-
task.add_done_callback(_on_done)
79-
80-
try:
81-
return await _asyncio.wait_for(new_fut, timeout)
82-
except _asyncio.TimeoutError:
83-
if not task.done():
84-
try:
85-
task.cancel()
86-
except Exception:
87-
_wait_for_logger.warning("Failed to cancel task", exc_info=True)
88-
raise
89-
90-
_xo.wait_for = _xo_wait_for
91-
92-
# Compatibility shim: older xoscar (<= 0.9.6) spawns each actor pool as
93-
# a daemon process, which makes the worker process unable to spawn the
94-
# model subprocess (Python raises ``AssertionError: daemonic processes
95-
# are not allowed to have children``). Patch multiprocessing so the
96-
# daemon flag is forced off when xoscar tries to start a child pool.
97-
import multiprocessing.process as _mp_process
98-
99-
if not getattr(_mp_process, "_xinf_daemon_patch_applied", False):
100-
_orig_start = _mp_process.BaseProcess.start
101-
102-
def _start_allow_daemon_children(self):
103-
# Pretend the current parent process is not a daemon so that
104-
# ``assert not _current_process._config.get('daemon')`` passes
105-
# inside multiprocessing.Process.start(). xoscar always launches
106-
# its worker pools as daemon=True; for xinference the worker
107-
# must spawn model inference subprocesses, so we relax this.
108-
cur = _mp_process._current_process
109-
saved = cur._config.get("daemon")
110-
cur._config["daemon"] = False
111-
try:
112-
return _orig_start(self)
113-
finally:
114-
if saved is None:
115-
cur._config.pop("daemon", None)
116-
else:
117-
cur._config["daemon"] = saved
118-
119-
_mp_process.BaseProcess.start = _start_allow_daemon_children
120-
_mp_process._xinf_daemon_patch_applied = True
121-
12240

12341
_install()
12442
del _install

xinference/core/worker.py

Lines changed: 3 additions & 22 deletions
Original file line numberDiff line numberDiff line change
@@ -946,28 +946,9 @@ async def _create_subpool(
946946
)
947947
env[env_name] = ",".join([str(dev) for dev in devices])
948948

949-
subpool_kwargs: Dict[str, Any] = {"env": env}
950-
if start_python is not None:
951-
subpool_kwargs["start_python"] = start_python
952-
try:
953-
subpool_address = await self._main_pool.append_sub_pool(**subpool_kwargs)
954-
except TypeError as e:
955-
# Older xoscar (<= 0.9.6) does not accept ``start_python``.
956-
# Drop it transparently so launch still works against an
957-
# environment that has not been upgraded yet.
958-
if "start_python" in str(e) and "start_python" in subpool_kwargs:
959-
logger.warning(
960-
"Installed xoscar does not support 'start_python' "
961-
"(requires >= 0.9.7). Falling back to the default "
962-
"Python interpreter; virtual environments configured "
963-
"via 'start_python' will not take effect."
964-
)
965-
subpool_kwargs.pop("start_python", None)
966-
subpool_address = await self._main_pool.append_sub_pool(
967-
**subpool_kwargs
968-
)
969-
else:
970-
raise
949+
subpool_address = await self._main_pool.append_sub_pool(
950+
env=env, start_python=start_python
951+
)
971952
await self._ensure_subpool_monitor()
972953
return subpool_address, [str(dev) for dev in devices]
973954

0 commit comments

Comments
 (0)