Peter Szemraj PRO
pszemraj
AI & ML interests
metallic intuition
Recent Activity
commentedon a paper about 15 hours ago
Puro-2B: Poor Lab's Qwen2-1.5B Trained on RTX 5090 within $5090 commentedon a paper 1 day ago
Puro-2B: Poor Lab's Qwen2-1.5B Trained on RTX 5090 within $5090 upvoted a paper 2 days ago
Puro-2B: Poor Lab's Qwen2-1.5B Trained on RTX 5090 within $5090Organizations
upvoted a paper 2 days ago
upvoted a paper 10 days ago
commented on FineBooks: are open OCR models good enough to unlock historical knowledge? 13 days ago
Hey, thanks for following through on that! I see the results on the leaderboard just now.
I need to sift through the different metrics to make sense of things because falconOCR is shockingly bad at least from the main ranking, but maybe it's an out-of-distribution thing, I'm not sure. Or maybe it's a case where it needs its little layout classifier model (or vice versa), which sometimes I find helps, and other times is terrible for quality.
- The being worse than Tesseract is interesting to say the least. Don't really have anything more to tell you yet..
One thing I did want to request, if possible, would be this new nemotron-parse 2.0 from Nvidia, which is very interesting simply from a license and origin perspective in that it is not a fine-tune of some Chinese VLM afaik, but a from-scratch architecture
upvoted an article 18 days ago
Article
Meet North Micro Vision: A 2.4B Native-Resolution Vision-Language Model
CohereLabs
β’ β’ 39Full-bandwidth transformer
Paper β’ 2608.08888 β’ Published β’ 23
Massive Activations in Hybrid Linear Attention Large Language Models: Pre-Attention Spikes and Inter-Spike Plateaus
Paper β’ 2608.12149 β’ Published β’ 30
LLMRouter: Unified Infrastructure for Developing, Evaluating, and Deploying LLM Routers
Paper β’ 2608.06867 β’ Published β’ 111
Poor Man's Agentic Modeling: Simulating Large LLM-Agent Societies on a Laptop
Paper β’ 2608.11215 β’ Published β’ 6