恒美微站
首页
关于我们
建站服务
主题模板
案例展示
资讯中心
联系我们
LM cache示例代码运行
首页
资讯中心
/
LM cache示例代码运行
LM cache示例代码运行
发布时间:2026/10/9 5:43:16
LMCache Installationhttps://docs.lmcache.ai/zh_CN/v0.3.6/getting_started/installation.html步骤1第一章 使用pip 安装步骤1UV 环境有一个比较合适的历史版本组合可以试vLLM0.10.0 LMCache0.3.3。**LMCache 仓库的 CacheBlend 问题记录里有人用这组版本https://github.com/LMCache/LMCache/issues/1823针对你要试的vLLM 0.10.0 LMCache 0.3.3建议用PyTorch 2.7.1 CUDA 12.8。这是有官方依据的LMCachev0.3.6 安装文档的兼容表明确标记vLLM 0.10.0.x LMCache 0.3.3兼容同页还写明 vLLM 0.10.0 使用 PyTorch 2.7.1。vLLM 0.10.0 的安装文档说明其预编译 CUDA 二进制基于 CUDA 12.8。LMCache v0.3.6 安装与兼容表 vLLM 0.10.0 安装文档uv venv--python3.12source.venv/bin/activate步骤2配置清华源mkdir-p~/.config/uvcat~/.config/uv/uv.tomlEOF [[index]] url https://mirrors.tuna.tsinghua.edu.cn/pypi/web/simple/ default true EOF步骤3克隆uv pipinstallvllm0.10.0--torch-backendcu128 uv pipinstalllmcache0.3.3步骤4模型下载python-mpipinstallmodelscope modelscope download\--modelAI-ModelScope/Mistral-7B-Instruct-v0.2\--local_dir/root/autodl-tmp/mfo/models/Mistral-7B-Instruct-v0.2步骤5修改gpu_worker.py参考\vllm\LMCache\examples\blend_kv_v1\README.md在已激活的(mfo)环境里运行python -c import vllm; print(vllm.__version__); print(vllm.__file__)(mfo) rootautodl-container-ujxcycmw77-dffbc8d4:~/autodl-tmp/mfo# python -c “import vllm; print(vllm.version); print(vllm.file)”0.10.0/root/autodl-tmp/mfo/.venv/lib/python3.12/site-packages/vllm/init.py(mfo) rootautodl-container-ujxcycmw77-dffbc8d4:~/autodl-tmp/mfo#步骤6示例执行##把实例的代码拷贝一下然后执行python blend.py\--model/root/autodl-tmp/mfo/models/Mistral-7B-Instruct-v0.2第二章 使用原始安装2.1 本地基于Lmcache 创建分支gitremote set-url official https://github.com/LMCache/LMCache.gitgitfetch official refs/tags/v0.3.3:refs/tags/v0.3.3gitswitch-cembed_url_033gitpush-uorigin embed_url_0332.2 服务器上拉去自己代码git checkout embed_url_033 #查询一下当前的版本 python -c import lmcache; print(lmcache.__file__) cd ~/autodl-tmp/mfo/LMCache # 确认当前用的是目标 uv 环境 which python python -m pip show lmcache # 安装当前源码不重新解析/替换环境里的依赖 uv pip install -e . --no-deps --no-build-isolation报错附录1transformers 版本过高≥5.0与 vllm 0.10.0 不兼容需要降级到 4.x 系列。uv pip install transformers4.55.2 python -c import transformers; print(transformers.__version__) python blend.py --model /root/autodl-tmp/mfo/models/Mistral-7B-Instruct-v0.22显存不足以容纳默认的 32K 上下文 KV cacheValueError: To serve at least one request with the models’s max seq len (32648), (3.99 GiB KV cache is needed, which is larger than the available KV cache memory (2.07 GiB). Based on the available memory, the estimated maximum model length is 16928. Try increasinggpu_memory_utilizationor decreasingmax_model_lenwhen initializing the engine.日志显示vLLM 可用 KV cache 显存为2.07 GiB但max_model_len32648需要3.99 GiB按当前可用显存最大上下文长度约为16928 tokens。在blend.py创建EngineArgs(...)的位置设置较小的max_model_len例如先试14000llm_args EngineArgs( modelmodel, max_model_len14000, # 其余参数保持不变 )如果脚本已有max_model_len改那个值即可。这样会降低单次请求允许的最大上下文长度模型本身不会改变。之后重新运行若你的测试 prompt 超过 14K tokens再逐步调高但不要超过日志估算的约 16.9K。