Unsloth Dynamic 3.0 GGUFs

(unsloth.ai)

63 points | by jonesy827 1 hour ago

8 comments

  • xlayn 50 minutes ago
    Hey Unsloth, your gguf are the first ones I look for when I want to download a gguf model. Today I was trying in fact to see, what's the smallest Qwen3.8-27B that I could run and get good results, say restricting it to 16GB of ram.. so I went, pick up the Qwen3.8-27B-UD-IQ2_XXS.gguf and them BAM, error on MTP... now I understand why after reading your announcement. Beyond the space saving, why removing the MTP? improves speed exactly for the group that could benefit from it.
    • mike-the-brain 49 minutes ago
      you can still have it, no?

      > We also removed the MTP module from smaller quants under UD-Q2_K_XL (8.37GB and lower) to converse around 500MB of disk space - you can use the Q4_0 MTP separate module if needed

      • xlayn 18 minutes ago
        my bad, you are totally right, thanks!
  • QuantumNomad_ 3 minutes ago
    Is it possible to use a model that needs around 64 GB VRAM if you have four GPUs with 16 GB VRAM each?
  • mike-the-brain 50 minutes ago
    Might be off-topic but: is it possible to perform such a quantization on Apple devices? Something like Mac Studio Ultra M1 (even if it would take weeks/months)?
  • tetsuo420 27 minutes ago
    It seems the NVFP4 quants have a preview version of this Unsloth Dynamic 3.0. Is this close to the finished version, or would it be better to switch to one of the newer quants?
  • throwa356262 47 minutes ago

       "We also made some smaller UD-1bit quants with UD-IQ1_S being 6.2GB (without MTP) which retain around 72% top-1% accuracy yet being 89% smaller"
    
    
    This is crazy! But has anyone tried these lower quants on real projects?
    • kennywinker 37 minutes ago
      Not 1-bit, but I’m getting pretty good results with some light coding using unsloth’s previous 2-bit quant of qwen3.8-27b. With these new quants i may be able to bump up to 3bit, tho it’s already running so slow (15tok/s average for the first 32k of context) that the speed hit might make it not worth the extra smarts
  • jadbox 41 minutes ago
    The new IQ4XS has been working pretty well so far on 4090 16gb.
    • kamranjon 26 minutes ago
      What size context are you able to squeeze in with less than 2gb of headroom? I have had some luck using a quantized kv cache but i fear that also decreases overall quality.
  • spwa4 24 minutes ago
    No MLX versions for 3.8 though.
  • lostmsu 13 minutes ago
    Cool. Now run TerminalHard and compare to unquantized 27B.

    KLD of 1%, or similar error metric that multiplies, on 10000 tokens would give accumulated error of 2,000,000%