Hugging Face
Models
Datasets
Spaces
Buckets
new
Docs
Enterprise
Pricing
Website
Tasks
HuggingChat
Collections
Languages
Organizations
Community
Blog
Posts
Daily Papers
Hardware
Learn
Discord
Forum
GitHub
Solutions
Team & Enterprise
Hugging Face PRO
Enterprise Support
Inference Providers
Inference Endpoints
Storage Buckets
Log In
Sign Up
🔄
In a Training Loop
102764.1
TFLOPS
Lewis Tunstall
PRO
lewtun
1261
261
899
Follow
Joflores's profile picture
omiom33's profile picture
roseking's profile picture
1,475 followers
·
138 following
https://lewtun.github.io/blog/
_lewtun
lewtun
AI & ML interests
LLMs, LLMs, LLMs
Recent Activity
updated
a collection
about 18 hours ago
MLE
updated
a collection
about 18 hours ago
MLE
updated
a bucket
3 days ago
lewtun/trl-internal-testing
View all activity
Organizations
lewtun
's activity
All
Models
Datasets
Spaces
Buckets
Papers
Collections
Community
Posts
Upvotes
Likes
Articles
New activity in
rl-llm-wiki/knowledge-base
about 1 month ago
source: arxiv:2603.16206 - OXA fine-tuning
2
#564 opened about 1 month ago by
lewtun
source: arxiv:2305.00944 - Poisoning instruction tuning
2
#562 opened about 1 month ago by
lewtun
New activity in
attention-wiki/knowledge-base
about 2 months ago
Process arXiv:2310.01889 - Ring Attention
5
#19 opened about 2 months ago by
lewtun
Process arXiv:2309.17453 - StreamingLLM
2
#10 opened about 2 months ago by
lewtun
Add source: Retrieval Head Mechanistically Explains Long-Context Factuality (arxiv:2404.15574)
1
#31 opened about 2 months ago by
lvwerra
Add source: H2O — Heavy-Hitter KV-cache eviction (arxiv:2306.14048)
2
#29 opened about 2 months ago by
lvwerra
Add source: NoPE — positional encoding & length generalization (arxiv:2305.19466)
2
#33 opened about 2 months ago by
lvwerra
Add sources: the 'attention as explanation' debate — Jain&Wallace + Wiegreffe&Pinter
2
#32 opened about 2 months ago by
lvwerra
Add source: In-context Learning and Induction Heads (arxiv:2209.11895)
2
#30 opened about 2 months ago by
lvwerra
Add sources: T5, DeBERTa, TUPE — relative & disentangled positional encoding
2
#26 opened about 2 months ago by
lvwerra
Add source: GQA — Grouped-Query Attention (arxiv:2305.13245)
4
#21 opened about 2 months ago by
lvwerra
Add source: Shaw et al. — Self-Attention with Relative Position Representations
2
#20 opened about 2 months ago by
lvwerra
Process arXiv:2309.06180 - PagedAttention
2
#9 opened about 2 months ago by
lewtun
Process arXiv:2309.00071 - YaRN
2
#8 opened about 2 months ago by
lewtun
Process arXiv:2307.03172 - Lost in the Middle
2
#7 opened about 2 months ago by
lewtun
Process arXiv:2306.15595 - Position Interpolation
2
#6 opened about 2 months ago by
lewtun
Process arXiv:2108.12409 - ALiBi
2
#5 opened about 2 months ago by
lewtun
Process arXiv:1911.02150 - Multi-query attention
2
#4 opened about 2 months ago by
lewtun
Process arXiv:1901.02860 - Transformer-XL
2
#3 opened about 2 months ago by
lewtun
Process arXiv:2104.09864 - RoFormer/RoPE
3
#2 opened about 2 months ago by
lewtun
Load more