<?xml version="1.0" encoding="utf-8" standalone="yes"?><rss version="2.0" xmlns:atom="http://www.w3.org/2005/Atom"><channel><title>模型推理 - 标签 - 罗孚传说</title><link>https://rovertang.com/tags/%E6%A8%A1%E5%9E%8B%E6%8E%A8%E7%90%86/</link><description>罗孚传说记录罗孚在 AI、本地模型、地图、汽车与数字办公中的实践。面对真实问题，分享动手过程、试错经验，以及对技术和工作的独立思考。</description><generator>Hugo 0.167.0 &amp; FixIt v1.0.0-alpha.3</generator><language>zh-CN</language><lastBuildDate>Sun, 25 May 2025 00:00:00 +0800</lastBuildDate><atom:link href="https://rovertang.com/tags/%E6%A8%A1%E5%9E%8B%E6%8E%A8%E7%90%86/index.xml" rel="self" type="application/rss+xml"/><item><title>SXM2版V100显卡很麻烦但很香：14B大模型速度超50！附折腾攻略</title><link>https://rovertang.com/llm/v100-sxm2-llm/</link><pubDate>Sun, 25 May 2025 00:00:00 +0800</pubDate><guid>https://rovertang.com/llm/v100-sxm2-llm/</guid><category domain="https://rovertang.com/categories/llm/">本地模型</category><description>SXM2版V100显卡非常具有性价比，适合AI大模型，单卡16GB版本一千出头，就能有不错的推理速度体验。但确实非常的折腾，转接卡、散热器、安装过程、系统兼容问题、驱动问题等都有不少的麻烦，坑有点多。只建议动手能力强的朋友入手。</description></item><item><title>大模型显卡推理和纯CPU推理对比测试</title><link>https://rovertang.com/llm/gpu-cpu-llm-speed/</link><pubDate>Sun, 04 May 2025 00:00:00 +0800</pubDate><guid>https://rovertang.com/llm/gpu-cpu-llm-speed/</guid><category domain="https://rovertang.com/categories/llm/">本地模型</category><description>通过256GB内存以及2张P106显卡下的多场景测试，证明了一个毋庸置疑的结论：显卡推理完胜纯CPU推理以及混合推理。虽然证明了一个寂寞，但也给低成本多显卡推理带来了一些希望。</description></item><item><title>两千元服务器跑671B大模型：能跑，看你想要什么。</title><link>https://rovertang.com/llm/cpu-671b-llm/</link><pubDate>Sat, 05 Apr 2025 00:00:00 +0800</pubDate><guid>https://rovertang.com/llm/cpu-671b-llm/</guid><category domain="https://rovertang.com/categories/llm/">本地模型</category><description>两千元服务器，也能运行671B DeepSeek-R1！虽速度不快，但性价比极高，是理(无)想(奈)的本地部署纯CPU推理方案。本文探讨了部署目的，回顾了翻车原因，进行了速度测试，提供了质量参考，还提出了后续方案，并附上模型文件与测试代码下载链接，欢迎大家沟通交流。</description></item><item><title>纯CPU推理大模型服务器翻车了</title><link>https://rovertang.com/llm/cpu-llm-server-failure/</link><pubDate>Sat, 08 Mar 2025 00:00:00 +0800</pubDate><guid>https://rovertang.com/llm/cpu-llm-server-failure/</guid><category domain="https://rovertang.com/categories/llm/">本地模型</category><description>花费一两千元使用E5 CPU搭建的纯CPU推理70B大模型服务器翻车了，CPU指令集和DDR4内存带宽是致命问题，以后对于不支持AVX512或AMX指令集的CPU还是不要考虑了吧。</description></item><item><title>Intel版MacBook用ollama跑大模型：显卡没用上，靠CPU硬扛。</title><link>https://rovertang.com/llm/macbook-ollama-llm/</link><pubDate>Wed, 12 Feb 2025 00:00:00 +0800</pubDate><guid>https://rovertang.com/llm/macbook-ollama-llm/</guid><category domain="https://rovertang.com/categories/llm/">本地模型</category><description>Intel版MacBook用ollama跑大模型，显卡没用上，只能靠CPU硬扛。如果改用llama.cpp等程序来跑大模型，应该可以将独立显卡利用起来。而通过输出token速度测试，侧面证明“使用CPU硬扛大模型”的方式也完全可行。</description></item><item><title>百元P106显卡跑7B大模型，矿渣变AI神器，真香！</title><link>https://rovertang.com/llm/p106-local-llm/</link><pubDate>Sat, 01 Feb 2025 00:00:00 +0800</pubDate><guid>https://rovertang.com/llm/p106-local-llm/</guid><category domain="https://rovertang.com/categories/llm/">本地模型</category><description>区区一百元，装块矿渣显卡就能让老电脑焕然一新，能玩黑神话，能跑大模型，不得不说P106显卡真香！</description></item></channel></rss>