AI & Algorithms
Attention Is All You Need:Transformer 架構完整拆解
拆解原始 Transformer 論文的 scaled dot-product attention、multi-head attention、encoder/decoder 區塊、殘差與 Layer Normalization、位置編碼與訓練方式,並區分 2017 年原始論文與現代 decoder-only LLM 在架構細節上的實際差異。
大语言模型的原理、训练与应用。
AI & Algorithms
拆解原始 Transformer 論文的 scaled dot-product attention、multi-head attention、encoder/decoder 區塊、殘差與 Layer Normalization、位置編碼與訓練方式,並區分 2017 年原始論文與現代 decoder-only LLM 在架構細節上的實際差異。