attention-mechanism.
Listed 12Updated Sep 7, 2026
- Rank 1. vit-pytorchImplementation of Vision Transformer, a simple way to achieve SOTA in vision classification with only a single transformer encoder, in Pytorchmit · linuxGitHub stars: 25.5k
- Rank 2. x-transformersA concise but complete full-attention transformer with a set of promising experimental features from various papersmit · linuxGitHub stars: 5.9k
- Rank 3. soundstorm-pytorchImplementation of SoundStorm, Efficient Parallel Audio Generation from Google Deepmind, in Pytorchmit · linuxGitHub stars: 1.5k
- Rank 4. perceiver-pytorchImplementation of Perceiver, General Perception with Iterative Attention, in Pytorchmit · linuxGitHub stars: 1.2k
- Rank 5. OpenSTLOpenSTL: A Comprehensive Benchmark of Spatio-Temporal Predictive Learningapache-2.0 · linuxGitHub stars: 1.1k
- Rank 6. tab-transformer-pytorchImplementation of TabTransformer, attention network for tabular data, in Pytorchmit · linuxGitHub stars: 1.1k
- Rank 7. enformer-pytorchImplementation of Enformer, Deepmind's attention network for predicting gene expression, in Pytorchmit · linuxGitHub stars: 573
- Rank 8. slot-attentionImplementation of Slot Attention from GoogleAImit · linuxGitHub stars: 498
- Rank 9. MultiModalMambaA novel implementation of fusing ViT with Mamba into a fast, agile, and high performance Multi-Modal Model. Powered by Zeta, the simplest AI framework ever.mit · linuxGitHub stars: 474
- Rank 10. deformable-attentionImplementation of Deformable Attention in Pytorch from the paper "Vision Transformer with Deformable Attention"mit · linuxGitHub stars: 368
- Rank 11. se3-transformer-pytorchImplementation of SE3-Transformers for Equivariant Self-Attention, in Pytorch. This specific repository is geared towards integration with eventual Alphafold2 replication.mit · linuxGitHub stars: 335
- Rank 12. bidirectional-cross-attentionA simple cross attention that updates both the source and target in one stepmit · linuxGitHub stars: 199