跳转至

参考文献汇总

更新日期:2026-04-15


一、基础架构论文

1.1 Transformer 核心

  • Vaswani et al. Attention Is All You Need. 2017. 论文

  • Radford et al. GPT-2. 2019. 论文

  • Brown et al. GPT-3. 2020. 论文

  • Touvron et al. LLaMA. 2023. 论文

  • Touvron et al. LLaMA-2. 2023. 论文

  • Meta. LLaMA-3 Technical Report. 2024. 论文

1.2 注意力变体

  • Shazeer. MQA: Fast Transformer Decoding. 2019. 论文

  • Ainslie et al. GQA. 2023. 论文

  • DeepSeek-AI. DeepSeek-V2 (MLA). 2024. 论文

  • DeepSeek-AI. DeepSeek-V3. 2024. 论文

  • Yang et al. GLA: Gated Linear Attention. 2023. 论文

1.3 位置编码

  • Su et al. RoPE. 2021. 论文

  • Press et al. ALiBi. 2021. 论文

  • Peng et al. YaRN. 2023. 论文

  • Ding et al. LongRoPE. 2024. 论文

  • Kazemnejad et al. Positional Encoding Impact. 2023. 论文

  • NoPE. 2024. 论文

  • Golovneva et al. CoPE. 2024. 论文

1.4 MoE

  • Shazeer et al. Sparsely-Gated MoE. 2017. 论文

  • Fedus et al. Switch Transformer. 2022. 论文

  • Jiang et al. Mixtral of Experts. 2024. 论文

  • Dai et al. DeepSeekMoE. 2024. 论文

  • Wang et al. Aux-Loss-Free Load Balancing. 2024. 论文

1.5 SSM / Mamba

  • Gu & Dao. Mamba. 2023. 论文

  • Dao & Gu. Mamba-2. 2024. 论文

  • Lieber et al. Jamba. 2024. 论文

  • Peng et al. RWKV. 2023. 论文


二、训练技术

2.1 优化器

  • Kingma & Ba. Adam. 2014. 论文

  • Loshchilov & Hutter. AdamW. 2017. 论文

  • Jordan. Muon Optimizer. 2024. 博客

  • Liu & Su. Muon is Scalable. 2025. 论文

  • NorMuon. 2025. 论文

  • Yang et al. μP. ICLR 2022. 论文

2.2 分布式训练

  • Shoeybi et al. Megatron-LM. 2019. 论文

  • Narayanan et al. Efficient Training on GPU Clusters. 2021. 论文

  • Rajbhandari et al. ZeRO. 2020. 论文

  • Huang et al. GPipe. 2019. 论文

  • Narayanan et al. PipeDream. 2019. 论文

  • Qi et al. Zero Bubble Pipeline. 2024. 论文

  • Liu et al. Ring Attention. 2023. 论文

  • Korthikanti et al. Reducing Activation Recomputation. 2022. 论文

2.3 Flash Attention

  • Dao et al. FlashAttention. NeurIPS 2022. 论文

  • Dao. FlashAttention-2. 2023. 论文

  • Shah et al. FlashAttention-3. 2024. 论文

2.4 Scaling Laws

  • Kaplan et al. Scaling Laws. 2020. 论文

  • Hoffmann et al. Chinchilla. 2022. 论文

  • Understanding LLM via Compression. 2025. 论文

  • Compression Laws for LLMs. 2025. 论文

2.5 FP8 训练

  • Micikevicius et al. FP8 Formats. 2022. 论文

  • Peng et al. FP8-LM. 2023. 论文


三、数据工程

  • Penedo et al. FineWeb. 2024. 论文

  • Gunasekar et al. Textbooks Are All You Need. 2023. 论文

  • Shumailov et al. The Curse of Recursion (Model Collapse). 2024. 论文

  • Wang et al. Self-Instruct. 2022. 论文

  • Xu et al. WizardLM/Evol-Instruct. 2023. 论文

  • Xu et al. Magpie. 2024. 论文

  • Demystifying Synthetic Data. 2025. 论文



上级 · Z3. 参考文献汇总