Qwen 3.8-Flash-Next releasing tomorrow (125B a6B)

(modelscope.cn)

93 points | by garo-pro 2 hours ago

16 comments

  • notnullorvoid 15 minutes ago
    It will be interesting to see the intersection of this with inference engines like FreeToken which improve distribution of work for MoE models across CPU/RAM and GPU/VRAM.

    If all it takes for a competitive model to run locally at good speeds is a used 3090 and some DDR4, then we might be in for the year of local AI.

    https://github.com/FlashML-org/FreeToken

    • Zylokloto 4 minutes ago
      You can already run it locally its just not the same.

      It is still slow, a lot slower than what you are used to with claude and co.

      And as soon as you increase context size, your memory requirements jump.

      Then when it runs for 30 minutes for something claude needs 5, your device will get hot.

      And even a used 3090 is apparently now between 1-2k.

  • ddtaylor 48 minutes ago
    I enjoy the Qwen models a lot, but building things on top of them with OpenRouter has been painful.

    OpenRouter does a lot of great work and I really enjoy being able to use different models so easily. I like when a provider is phasing out an older model that still works for my needs and the price is much lower. It seems like such a good win-win.

    However, the problem is that many Qwen models have almost no capacity or is so flaky you literally have to just litter your code with a blacklist/whitelist of providers. OpenRouter has some attempts to solve this, but they don't work. In fact, OpenRouter has a lot of really cool stuff that is documented, but if you read the code it's not yet implemented or isn't actually there yet, which is a shame.

    I tried to get in contact with them at OpenRouter about this and I was interested in working with them in the past, but it's difficult to get in touch with the right people and they are growing very fast. I expect being acquired by Stripe will accelerate those problems in some ways. I have no doubt they will resolve all of these issues eventually and scaling that much that quickly is really hard, so kudos to them, but the road has been pretty lame and taken some wind out of my sails.

    • geek_at 18 minutes ago
      The best solution to this for me is to self host litellm or a different router and use model aliases. For example I have a model called "coding" and when a new good model comes out I just switch the backend without needing to change the alias or the key in my projects (opencode, etc).

      I have a few of them even a smart router called "agents" which will use local models but if it thinks the request might require higher reasoning it's routing to a different model

      • try-working 6 minutes ago
        I built a router that lets you route between local and cloud models. Link in my profile.
    • irthomasthomas 35 minutes ago
      Openrouter was pretty great before prompt caching became common. Now it is extremely expensive for most individual workflows, unless you spend a lot of work customizing router preferences, and then you still get a worse cache hit rate than using the provider directly. I only keep $5-$10 in OR for occasional testing.
    • npn 6 minutes ago
      I'm confused? Can you just define some presets and call them instead? With preset you can pinpoint a lot of things, especially the providers
  • fcanesin 1 hour ago
  • pwython 53 minutes ago
    I was already rolling around the idea of a 128GB M5 Max MBP. Now this!

    A 4-bit MLX quant with 128k window should fit perfectly, in the 50-70 tok/s range.

    • irthomasthomas 30 minutes ago
      IDK, prefill speed is a bigger concern for most wokflows, like agent coding, and I heard that this is quite low on macs?
      • smcleod 27 minutes ago
        That was mainly before the M4 generation when they didn't have matmul instructions.
        • jasonjmcghee 19 minutes ago
          M5 prefill is much faster than M4.

          I've seen benchmarks that show 4-5x faster of M5 Max vs. M4 Max.

          For local models you're likely using M5 Max, prefill is low thousands of tokens per second, as opposed to, say high hundreds with M4 Max.

          For larger dense models, some fraction of that, but similar multiple.

          • smcleod 8 minutes ago
            Yes, I have the M5 Max. But there was no matmul acceleration before the M4 which made things a lot slower.
    • sscaryterry 50 minutes ago
      I have a 128GB M5 Max, and it sucks at this stage. 50-70 tok/s might be something...
      • smcleod 25 minutes ago
        50-70tk/s is what I get on my m5 max on a 5-6bit Qwen 3.8 27B?
  • hedora 16 minutes ago
    Time to dust off my 128GB strix halo (literally—it’s been dusty, and it’s running a bit warm these days).

    Any idea where this model sits according toquality benchmarks? Pre-bubble MSRP on this hardware was $1400, and it draws 200-ish watts, putting it down into consumer territory.

    I’m wondering if it can replace claude for llm-friendly coding tasks.

  • big-chungus4 53 minutes ago
    > We are releasing these architectural improvements ahead of time so that the community can prepare for the upcoming full family of Qwen4 models.

    That gives me hope that "full family" means it will include smaller models like 4B.

  • honestlyranked 1 hour ago
    Alibaba is giving sleepless nights to the tech giants
  • big-chungus4 58 minutes ago
    I hope there is going to be a free endpoint... Unlike 35B-A3B, I am nowhere close to running it locally
  • cogman10 1 hour ago
    Wow. I wasn't expecting this. I thought they were going to do a 35B model instead.
    • hasteg 21 minutes ago
      As a 5090 owner and local model enthusiast, I was hoping it would be 35B A3B so I could run it myself =(.
      • cpburns2009 6 minutes ago
        You can run the 27B released last week. I haven't tried it yet myself but the 3.6 version runs great on my 5090.
  • BrucecarlL 47 minutes ago
    Waiting for the performance report! Ai hope it can beat DS
  • tarruda 2 hours ago
    Can you share the source for the parameter count (125B A6B)? I didn't see it anywhere in the page.
  • bellowsgulch 40 minutes ago
    Really happy for those with 128GB+ RAM. Sitting here with my Apple M1 Max with 64GB though. Was looking forward to a Qwen3.8-35B-A3B like many others.
    • dofm 21 minutes ago
      Have you tested Muse Glimmer in low reasoning strength?

      Token generation is slow (and prefill is) but you will likely find it solves actual problems faster than Qwen 3.6 35B-A3B.

  • tw1984 52 minutes ago
    Qwen4 sounds exciting
  • blurbleblurble 57 minutes ago
    gg
    • david927 10 minutes ago
      Well put and succinctly put. And if OxA is a flash model? it becomes: goodnight
  • mrdoe 1 hour ago
    lol blocked with dns4eu

    what a joke this resolver has become

  • metrofun 24 minutes ago
    [dead]