Skip to content

attention-mechanisms.

Listed 17Updated Aug 28, 2026

  1. Rank 1. PaLM-rlhf-pytorchImplementation of RLHF (Reinforcement Learning with Human Feedback) on top of the PaLM architecture. Basically ChatGPT but with PaLMGitHub stars: 7.9k
  2. Rank 2. audiolm-pytorchImplementation of AudioLM, a SOTA Language Modeling Approach to Audio Generation out of Google Research, in PytorchGitHub stars: 2.6k
  3. Rank 3. alphafold3-pytorchImplementation of Alphafold 3 from Google Deepmind in PytorchGitHub stars: 1.7k
  4. Rank 4. BS-RoFormerImplementation of Band Split Roformer, SOTA Attention network for music source separation out of ByteDance AI LabsGitHub stars: 935
  5. Rank 5. magvit2-pytorchImplementation of MagViT2 Tokenizer in PytorchGitHub stars: 668
  6. Rank 6. mmditImplementation of a single layer of the MMDiT, proposed in Stable Diffusion 3, in PytorchGitHub stars: 559
  7. Rank 7. iTransformerUnofficial implementation of iTransformer - SOTA Time Series Forecasting using Attention networks, out of Tsinghua / Ant groupGitHub stars: 541
  8. Rank 8. local-attentionAn implementation of local windowed attention for language modelingGitHub stars: 505
  9. Rank 9. recurrent-memory-transformer-pytorchImplementation of Recurrent Memory Transformer, Neurips 2022 paper, in PytorchGitHub stars: 425
  10. Rank 10. q-transformerImplementation of Q-Transformer, Scalable Offline Reinforcement Learning via Autoregressive Q-Functions, out of Google DeepmindGitHub stars: 406
  11. Rank 11. clinical-calculator-tooluseExplorations into training LLMs to use clinical calculators from patient history, using open sourced models. Will start with Wells' CriteriaGitHub stars: 316
  12. Rank 12. CoLT5-attentionImplementation of the conditionally routed attention in the CoLT5 architecture, in PytorchGitHub stars: 231
  13. Rank 13. simple-hierarchical-transformerExperiments around a simple idea for inducing multiple hierarchical predictive model within a GPTGitHub stars: 228
  14. Rank 14. MambaTransformerIntegrating Mamba/SSMs with Transformer for Enhanced Long Context and High-Quality Sequence ModelingGitHub stars: 228
  15. Rank 15. flash-cosine-sim-attentionImplementation of fused cosine similarity attention in the same style as Flash AttentionGitHub stars: 221
  16. Rank 16. recurrent-interface-network-pytorchImplementation of Recurrent Interface Network (RIN), for highly efficient generation of images and video without cascading networks, in PytorchGitHub stars: 210
  17. Rank 17. coconut-pytorchImplementation of 🥥 Coconut, Chain of Continuous Thought, in PytorchGitHub stars: 184