Hugging Face
Models
Datasets
Spaces
Buckets
new
Docs
Enterprise
Pricing
Website
Tasks
HuggingChat
Collections
Languages
Organizations
Community
Blog
Posts
Daily Papers
Hardware
Learn
Discord
Forum
GitHub
Solutions
Team & Enterprise
Hugging Face PRO
Enterprise Support
Inference Providers
Inference Endpoints
Storage Buckets
Log In
Sign Up
132.4
TFLOPS
Rider Jones
mazuj2
5
2
5
Follow
21world's profile picture
1 follower
·
12 following
AI & ML interests
None yet
Recent Activity
reacted
to
FredyRivera-dev
's
post
with ❤️
11 days ago
We wrote a full technical guide on how to train a bilingual (ES/EN) LLM from scratch: TinyQwen. Covers: - Hybrid architecture based on Qwen3.5 - Pre-training with 15B tokens - Cost benchmark between H200 and B200 - Post-training with SFT + LoRA - Full code and data, open source With ~$11 of compute on an H200 we ran an initial training run, enough to validate the full architecture and pipeline. Blog post: https://aquiles-ai.vercel.app/blog/tinyqwen-from-scratch Technical feedback welcome, especially from anyone looking to replicate the pipeline with more compute.
liked
a model
14 days ago
mindlab-research/Macaron-V1-Tall
new
activity
about 1 month ago
unsloth/Qwen3.6-27B-MTP-GGUF:
FAST!!!! 39tps!
View all activity
Organizations
None yet
mazuj2
's activity
All
Models
Datasets
Spaces
Buckets
Papers
Collections
Community
Posts
Upvotes
Likes
Articles
liked
a model
14 days ago
mindlab-research/Macaron-V1-Tall
Text Generation
•
36B
•
Updated
16 days ago
•
1.43k
•
68
liked
a model
5 months ago
turboderp/Qwen3.5-35B-A3B-exl3
Updated
Mar 2
•
255
•
22
liked
3 models
6 months ago
turboderp/Qwen3-VL-32B-Instruct-exl3
Updated
Nov 9, 2025
•
8
•
7
turboderp/Qwen3-Next-80B-A3B-Instruct-exl3
Updated
Nov 1, 2025
•
4
•
27
turboderp/Qwen3-VL-30B-A3B-Instruct-exl3
Updated
Nov 9, 2025
•
16
•
5