<?xml version="1.0" encoding="UTF-8"?><rss version="2.0" xmlns:atom="http://www.w3.org/2005/Atom"><channel><title>MindlyAI’s Blog</title><description>不空谈概念，交付能上线的 AI。10 年开发经验，贯通后端、大数据、大模型，专注 AI 大模型应用开发：RAG / GraphRAG、Agent 工作流、Text2SQL、微调与评测。</description><link>https://llmnex.top</link><language>zh-CN</language><lastBuildDate>Wed, 23 Sep 2026 07:55:20 GMT</lastBuildDate><atom:link href="https://llmnex.top/rss.xml" rel="self" type="application/rss+xml"/><item><title>大模型分布式训练算法：五种并行策略与代价</title><link>https://llmnex.top/distributed-training-algorithms</link><guid isPermaLink="true">https://llmnex.top/distributed-training-algorithms</guid><description>数据并行、流水线并行、张量并行、专家并行与 ZeRO——逐个说清楚切的是什么、代价落在哪，以及面试里该怎么说。</description><pubDate>Wed, 23 Sep 2026 09:00:00 GMT</pubDate><category>分布式训练</category><category>分布式训练</category><category>数据并行</category><category>流水线并行</category><category>张量并行</category><category>ZeRO</category></item><item><title>ZeRO-3 核心流程：把参数也切开的那一档</title><link>https://llmnex.top/zero-3-core-flow</link><guid isPermaLink="true">https://llmnex.top/zero-3-core-flow</guid><description>参数、梯度、优化器状态全部按卡数切开，每卡只留 1/N；哪一层要用就临时把参数拼回来，算完立刻释放。</description><pubDate>Wed, 23 Sep 2026 08:50:00 GMT</pubDate><category>分布式训练</category><category>ZeRO</category><category>分布式训练</category><category>显存优化</category><category>数据并行</category></item><item><title>Ring AllReduce 核心流程：梯度是怎么聚合的</title><link>https://llmnex.top/ring-allreduce-core-flow</link><guid isPermaLink="true">https://llmnex.top/ring-allreduce-core-flow</guid><description>梯度切成 N 块、N 张卡排成环，两轮各走 N−1 步：每卡只搬约 2 份梯度大小的数据，就能拿到全量梯度之和。</description><pubDate>Wed, 23 Sep 2026 08:40:00 GMT</pubDate><category>分布式训练</category><category>Ring AllReduce</category><category>集合通信</category><category>数据并行</category><category>分布式训练</category></item><item><title>LoRA 参数高效微调：低秩分解与关键配置</title><link>https://llmnex.top/lora</link><guid isPermaLink="true">https://llmnex.top/lora</guid><description>冻结原权重，只训练两个小矩阵 A(d×r) 与 B(r×k) 去逼近 ΔW，可训练参数从 dk 降到 r(d+k)：挂在哪、秩 r 怎么选。</description><pubDate>Wed, 23 Sep 2026 08:30:00 GMT</pubDate><category>参数高效微调</category><category>LoRA</category><category>PEFT</category><category>低秩分解</category><category>微调</category></item><item><title>QLoRA 原理图解：三个组件把微调显存压下来</title><link>https://llmnex.top/qlora</link><guid isPermaLink="true">https://llmnex.top/qlora</guid><description>4-bit 量化冻结基座 + 16-bit LoRA 适配器，按「传统量化 → NF4 → 双重量化 → 分页优化器」的顺序逐层拆开。</description><pubDate>Wed, 23 Sep 2026 08:20:00 GMT</pubDate><category>参数高效微调</category><category>QLoRA</category><category>量化</category><category>NF4</category><category>PEFT</category><category>显存优化</category></item><item><title>位置编码：让注意力理解 Token 在序列中的位置</title><link>https://llmnex.top/positional-encoding</link><guid isPermaLink="true">https://llmnex.top/positional-encoding</guid><description>从 X = E + P 走到 RoPE：为什么 Self-Attention 天生看不见顺序，位置信息又如何只靠点积留下相对距离。</description><pubDate>Wed, 23 Sep 2026 08:10:00 GMT</pubDate><category>模型原理</category><category>位置编码</category><category>RoPE</category><category>Transformer</category><category>Attention</category></item><item><title>激活函数：为神经网络注入非线性</title><link>https://llmnex.top/activation-functions</link><guid isPermaLink="true">https://llmnex.top/activation-functions</guid><description>从 Sigmoid、Tanh 到 ReLU 家族与 GELU / SwiGLU：梯度消失、神经元死亡，以及 FFN 里到底该用哪个。</description><pubDate>Wed, 23 Sep 2026 08:00:00 GMT</pubDate><category>模型原理</category><category>激活函数</category><category>ReLU</category><category>GELU</category><category>SwiGLU</category><category>FFN</category></item><item><title>前向传播与反向传播概述</title><link>https://llmnex.top/forward-and-backward</link><guid isPermaLink="true">https://llmnex.top/forward-and-backward</guid><description>微调场景下的前向产物、反向传播机制，以及冻结参数与 LoRA 等不同策略对反向计算图的实际影响。</description><pubDate>Wed, 23 Sep 2026 07:50:00 GMT</pubDate><category>模型原理</category><category>前向传播</category><category>反向传播</category><category>计算图</category><category>微调</category></item><item><title>梯度下降与优化器：从 SGD 到 Muon</title><link>https://llmnex.top/gradient-descent-and-optimizers</link><guid isPermaLink="true">https://llmnex.top/gradient-descent-and-optimizers</guid><description>把 θ ← θ − η∇L(θ) 这行公式拆开，看它落到工程里会分出哪四个问题，再按时间顺序看每一代优化器回答了哪一个。</description><pubDate>Wed, 23 Sep 2026 07:40:00 GMT</pubDate><category>模型原理</category><category>梯度下降</category><category>SGD</category><category>Adam</category><category>AdamW</category><category>Muon</category></item></channel></rss>