What inference engine/runtime is needed for this to be run?
#2
by Smoffyy - opened
I'm new to all these new Ternary models, and I'm wondering what engine is compatible for inference.
i had qwen remake their webgpu inference engine here is the dl: "gofile.io/d/6qzoT9"
run two terminal from the main window for these commands:
python -m http.server 8080
python serve_local.py
then open the browser at localhost:8080/demo.html
Smoffyy changed discussion status to closed
Thank you!
Smoffyy changed discussion status to open
Smoffyy changed discussion status to closed
Hey everyone, thanks so much for your patience. The production model is up, and the GGUFs are up. llama.cpp fork: https://github.com/SyzygyResearch/llama.cpp-mach1