Hugging Face's logo Hugging Face
  • Models
  • Datasets
  • Spaces
  • Buckets new
  • Docs
  • Enterprise
  • Pricing
    • Website
      • Tasks
      • HuggingChat
      • Collections
      • Languages
      • Organizations
    • Community
      • Blog
      • Posts
      • Daily Papers
      • Hardware
      • Learn
      • Discord
      • Forum
      • GitHub
    • Solutions
      • Team & Enterprise
      • Hugging Face PRO
      • Enterprise Support
      • Inference Providers
      • Inference Endpoints
      • Storage Buckets

  • Log In
  • Sign Up
Ofir Zafrir's picture

Ofir Zafrir

ofirzaf
3 14 6
guybd's profile picture 21world's profile picture johnbox86's profile picture
·
  • ZafrirOfir
  • ofirzaf

AI & ML interests

Sparsity, Qunatization, Model Compression

Recent Activity

liked a model about 1 month ago
z-lab/Qwen3.6-35B-A3B-DFlash
upvoted an article about 1 month ago
Intel XPU Kernel Skill: LLM-driven Triton kernel optimization for the Hugging Face Kernel Hub
upvoted an article 5 months ago
Getting More from Your Test-Time Compute Budget with Portfolio Beam Search
View all activity

Organizations

Intel's profile picture Intel Labs's profile picture Need4Speed's profile picture il-eai-nlp's profile picture

authored 2 papers over 1 year ago

Q8BERT: Quantized 8Bit BERT

Paper • 1910.06188 • Published Oct 14, 2019 • 2

FastDraft: How to Train Your Draft

Paper • 2411.11055 • Published Nov 17, 2024 • 11
authored a paper about 3 years ago

An Efficient Sparse Inference Software Accelerator for Transformer-based Language Models on CPUs

Paper • 2306.16601 • Published Jun 28, 2023 • 4
Company
TOS Privacy About Careers
Website
Models Datasets Spaces Pricing Docs