Meta provider

Llama 3.1 8B

Meta's small workhorse. Runs comfortably on any 8GB+ card at Q4 or any 12GB+ at FP16.

Parameters
8B
Family
Meta
Context
128K tokens
FP16 weights
16GB
// where you can run it
llama3.1:8bLM StudiovLLMMLXoMLX
// hugging face stats (cached daily)
10.9M downloads · 6K likes · license: llama3.1 · updated 1 year ago

What you need to run this.

$ ./vrambudget --model llama-3-1-8b --by quant

// budgets shown at ctx 8K, concurrency 1, 15% safety headroom. Tune in the calculator →

Alternatives at this size.

$ grep --params similar catalog.json

Discussion.

$ gh discussion list

// sign in with github to leave a comment. threads live in the repo's discussions tab.