How difficult is distilling?

I remember a year or so ago when DeepSeek R1 came out and it was pretty quickly distilled into Llama 3 8b and Qwen 2.5 (?) 7b. Why don’t we see more distilled models? How expensive is it? How many tokens or prompts does it take?

submitted by /u/GreedyWorking1499
[link] [comments]

Leave a Comment

Your email address will not be published. Required fields are marked *

Scroll to Top