33 comments

  • miellaby 2 hours ago
    The paper explains absolutely everything as if it was a tutorial "how to made your own modern agentic LLM". They even tell how they made their dataset. https://aleph-alpha.com/downloads/tech-report.pdf ; It's the first time I see this level of openness.
    • zelphirkalt 2 hours ago
      In my opinion not being open about which data is ingested and trained on, and trying to make that a repeatable thing for a third party, is not worth being called "open". Glad they did that.
    • kingcauchy 30 minutes ago
      Yeah the pdf alone is awesome as a learning tool.
    • p-e-w 1 hour ago
      This alone makes it much more valuable than many high-profile releases despite not quite performing at the same level.
    • ofjcihen 1 hour ago
      Hopefully this becomes the new standard.

      It’s seemed crazy to me that anyone thought these could stay closed or even SHOULD be closed source.

      • Loquebantur 58 minutes ago
        Be cautious what you wish for. Tools don't tell you what to do with them.

        Open source LLMs "democratize" access to the "intelligence booster" that is AI. But while that has several benefits, it also has several downsides.

        Humanity has the serious problem of being underdeveloped in the "spiritual" department. Ethics is often considered some sort of lifestyle choice, but it's actually the difference between order and chaos in a society.

        Everybody being able to do anything means somebody will be able to do something you don't like. At an arbitrary scale.

        • Zigurd 27 minutes ago
          The main thing that bugs me about the risk's discussion around AI is the lack of specificity. Commenters here have a good grasp of the risks around finding vulnerabilities faster than they can be patched. That's good and it matches the applicability of LLMs to coding.

          But the applicability and the ROI of LLMs for other use cases than coding is a lot squishier. Also correspondingly the risks are unspecific.

          As for what to do, ethical disclosure of vulnerabilities provided a good framework for disclosing software vulnerabilities discovered with the assistance of lLMs. What is going to be novel and calls for our spiritual development in other domains?

          • Loquebantur 1 minute ago
            You're right, people weirdly lack imagination on what "higher intelligence" (minus ethics) actually affords you, let's have a look:

            What do average people currently want? They're taught, the most important thing was being rich. So they will ask their AI to make them rich. Most real life ways to get there are "sketchy" to say the least, usually downright unethical and anti-social, but US society turns a blind eye when the "Wolf of Wall Street" comes out on top and the schemes don't easily fit into average people's abilities of moral judgement.

            -> Large parts of US society suddenly engaging in all kinds of "semi-legal/hyper-illegal" fraud schemes, at the expense of already saturated environmental and societal resilience. Guaranteed collapse.

            Or, let's get rid of those pesky neighbors/wrong-colored people/annoying opinions? Again, "legal" is a pretty squishy concept and only really applies when you don't have the legal expertise to get around it. Now you can.

            Or, look at the basics: what is "real"? You only "know" because you trust certain people and institutions. Generative AI can help with that /s.

            It's not only about "building weapons of mass destruction". It's about doing the same shit as usual, but a thousand times faster/amplified. Look up poly-/metacrisis for starters. Going faster with AI when there's a wall in front of you isn't the best idea.

        • adrianN 19 minutes ago
          Of course openness is only worse than leaving everything under the control of a select cabal of you believe that cabal to be more ethical than the rest of us.
        • computerdork 48 minutes ago
          Ah, didn't think of this. Some rogue militia group might try to use this LLM (or create their own LLM based on this work) to help them create biological weapons or to do a mass hacking the infrastructure of targeted country.

          Wonder what safeguards Kolibri uses to prevent this? Or if they even can

          • Zigurd 24 minutes ago
            Chemical and biological warfare is hard. A cult in Japan created a mass casualty event using nerve gas. Which is the only somewhat "successful" terrorist WMD attack I know of. Knowledge of how to create these weapons isn't new. Guns and bombs are the most widely used terror weapons for a reason.
            • UberFly 5 minutes ago
              Enter LLMs to help through all those pesky hard parts.
        • ofjcihen 7 minutes ago
          Oh definitely, and I’m in the cybersecurity space so I’m already on the “worst case scenario committee” hah.

          But the alternative just seems… so much worse to me?

          A select few groups gating access to the ability to do everything seems like neo-fuedalism in the making.

          And to be fair even the gating that we do have (daybreak, CVP, etc.) is already being circumvented via keys being stolen and sold on the dark web.

  • tomComb 1 hour ago
    For a post to make such a big deal about sovereignty it is a bit misleading to not mention that the company is slated to be merged with Cohere, a Canadian company.

    And that is a good thing - no need to hide it. Given the growing cost of keeping up, these few non-US, non-Chinese companies really need to do more sharing of efforts and costs.

    Canada too is very much in need of sovereign AI options, but funding that on its own would be pretty much a waste of money. Would love to see this new German Canadian company cooperate with Mistral too, or maybe one of the Korean AI companies.

    • Loquebantur 1 hour ago
      Given that "sovereign" these days implicitly means independence from USA and China, and Canada being spiritually in the same boat as the EU, it's pretty spot on?

      The costs of "keeping up" aren't really growing, on the contrary.

  • niemandhier 3 hours ago
    I think at the moment the main thing a sovereign AI model needs to be good at is auditing the results of other models.

    Right now one could run an open model for most government applications and it would be good enough, you just cannot trust any of these.

    So having a sovereign controlled model audit the first one would basically act like a “trust adapter”.

    If the second model is cheap and fast enough, there is a business model.

    You don’t even need to audit all the intermediate steps, just tool calls and end results.

    • holaysuns 3 hours ago
      The issue with 'sovereign' is not about sneaky things a foreign entity might put in there, it's about control of the stack.

      Since no individual buyer cares that much about geostrategic issues, it's almost impossible to get some random European company to think about buying anything other than 'whatever US or China' are making.

      There has to be very concerted push to make a difference.

      Legal mandates around sovereignty (be very careful here) could make a difference.

      But there will also be market reaction: US/China companies will bend on some level and provide things like 100% EU hosted and even comply with some data source stuff.

      The only real path is to be more competitive as a continent.

      • Loquebantur 1 hour ago
        The EU AI act isn't primarily about "geostrategy". It's about protecting society.

        AI models act as a force multiplier for intelligence, in particular for generating information according to someone's wishes.

        I.e. deepfakes and social media mis-/disinformation campaigns are a thing and having powerful AI allows you to do those at scales that can overwhelm society's resilience.

        In general, even if you have "aligned" AI: aligned with whom or what?

        Whom are you comfortable with lording as a some demi-god over you, dictating what to believe?

        • holaysuns 10 minutes ago
          European regulations are 100% geostrategy and they are defending against external interests.

          Those laws would not exist if European champions were leading the world.

          "are you comfortable with lording as a some demi-god over you, dictating what to believe?"

          Yes, Europe handed over all of the decisions about everything to foreign powers, now they have to enact regulations to try to constrain it.

          Zuck et. al. make the investments decisions for Europe, by virtue of you all giving him the money and power to do that.

          Stop giving him the money/power, then this regulation won't exist

      • Oras 3 hours ago
        > Since no individual buyer cares that much about geostrategic issues

        You clearly haven’t worked in enterprise in the EU. Location of data processor is the first thing they check. That’s why every major cloud has regions with different offerings, not just for HA and redundancy

        • holaysuns 3 hours ago
          They care because the are 'required to', otherwise those 'offerings' wouldn't even exist. It's mostly the result of regulatory action.
          • JaggerJo 2 hours ago
            I’m working with clients (big companies) that care a lot - but not because they have to. They care because they lost trust in non EU players.
          • rokkamokka 1 hour ago
            They also care because it's becoming increasingly risky to rely on US companies given the current political situation
          • bdangubic 2 hours ago
            nope, that no longer holds any water. was true though couple of years ago…
            • holaysuns 15 minutes ago
              The argument holds most of the water.

              Yes - security concerns are now very real, but do not fall for this idea that anyone really cares about structural concerns.

              Tons of EU companies are still selling crap to Russia, happy to look the other way wile the bear devours a neighbour, as long as profits are there.

              What has happened is that Trump has given a face to the reality, and so companies are adjusting on some level.

              But 1) it will be nominal 2) Trump will be gone and the impetus will fade 3) the conglomerates will react 4) modest legislation will mean ...

              The can will get kicked down the road.

              The invasion of E. Europe by Russia has not even caused defence spending or the size of Armies to change, doctrine is barely changing.

              Nothing will make those beheamoths change other than other forms of structural concern.

              The 'Riet of the Right' - partly due to collapse of Auto / China (among many other things) might cause some change. But the change won't happen until well after the damage has been done.

              Brexit should have been an opportunity for institutional reform of the EU, but no - they blamed it all on populism.

              Certain governing entities in Europe will try to move away from Microsoft, they will be pulled back.

              There is hope, and certain champions can rise and causes pieces to collapse.

              If SUSE had any true entrepreneurial whereiwthal - they would create a true consumer / prosumer / enterprise-user friendly variation and brand their flavour of Unix - and make sure that all Euopean governments us it exclusively, which would cascade into widespread use.

              There are variations of insta, youtube, netflix that should all be Europe based that could 'theoretically happen' but it needs some structural impetus.

              Things usually don't change. Usually there needs to be a collapse and re-order of the system for that to happne.

              The people in charge just want to keep their jobs, their very high salaries, and protect their retirement, and will be happy to 'sell out' whatever other imperatives along the way.

              Ending with a positive not - I would say current conditions mean the change is now 'plausible, it not likely' whereas before it was 'not very plausible'. So there's a candle, a bit of wind, but not a lot of tinder or dry wood.

    • adishiktensward 29 minutes ago
      It depends on the use case, but I believe SLMs can do great with smaller tasks. People got so attached to 'general purpose' that they forgot software can be designed for specific, smaller use cases.
    • mawadev 2 hours ago
      That is a very good idea and there is a market for that, especially if a company manages to source the hardware and then install the stack on a german or EU customer site
  • spijdar 3 hours ago
    The absence of any comparison to Qwen3.8 Flash, another MoE model with a small-ish (6B) number of active parameters, is pretty striking. Instead, it's compared with Qwen3-Next 80B-A3B, a model released almost a full year ago.

    I get that doesn't invalidate the real "point" of the model, but...

    • bitexploder 22 minutes ago
      That is what stood out to me as well. Qwen 3.8 Flash can run on very limited hardware as well and it is at least as good as Sonnet 5 in benchmarks like DeepSWE. With its n-gram design you can get flash next running on very limited GPU resources, as little as 16GB of VRAM.

      People follow the latest frontier lab models with great attention and migrate to the next big model on their subscriptions. Meanwhile these local models have quietly gotten REALLY good. It is not even an exaggeration. It has happened in the last couple of months.

      "Local model you can run at 40 t/s on a gaming machine that is better than Opus 4.6" is way less exciting than "OpenAI IS DOING CRIME!!! OpenAI SOLVED NAVIER STOKES. DARIO SAYS GLM 5.3 BAD! SLOW DOWN THE FRONTIER!".

      (edit: also... totally ignore that 27B dense column over there where Qwen 3.8 27B beats Kolibri on nearly every single benchmark. Why would I choose to run this model?)

    • okamiueru 47 minutes ago
      Can't you infer the comparisons you would like from baseline results provided? There are better results in the sibling post on HN: https://aleph-alpha.com/en/blog/kolibri-has-landed-a-soverei...
      • spijdar 18 minutes ago
        Maybe?

        Qwen3.8 27B scored notably higher in most of the provided benchmarks, including the German-specific ones. The only "downside" is that inference is much more costly and slow, since it's a dense model.

        Qwen3.8 Flash-Next appears to usually "benchmark higher" than 27B, while remaining fast.

        I'm sure I could dig up the equivalent benchmarks for Flash and do the comparison myself, but as far as inference goes, it's messy. Consider that Qwen3.5 35B-A3B scores higher than Qwen3.6 on some of the German-specific benchmarks.

        So it seems superficially plausible that Qwen3.8 Flash-Next might not be "27B but faster" in the ways that are important for this model. Or it could just "be superior" in all ways.

        Either way, I don't think an LLM has to be "the best" at anything to be worthwhile, necessarily. And I kind of distrust benchmarks on top of that, so...

  • 9dev 4 hours ago
    Aleph Alpha is just a sad joke by now. The talent isn't there anymore, they never managed to catch up to the other labs, failed to deliver on several projects, and by now are just a cash grab for the investors.
    • gchamonlive 3 hours ago
      Only if you think in terms of short term gains, but thinking in the long run, doesn't really matter if these models are crap today, all models will eventually be obsolete, unless we plateau hard on every aspect of the tech.

      What matters is to have good sovereign models in 5, 10, 20 years time.

      • roncesvalles 3 hours ago
        I'm curious why you think sovereign models matter that much if open source ones exist (unless that stops at some point). Sovereign model hosting services, yes.
        • HotHotLava 3 hours ago
          Just open weights are not enough, I think as a nation state you'd at least want a whole sovereign training pipeline. Otherwise you are stuck when Kimi K4 stops releasing their weights.

          But I don't see why this needs to be done on a state-by-state level, a sovereign EU model seems to be much more realistic in terms of funding.

          • rixed 2 hours ago
            Wouldn't you also want to be able to manufacture some sort of GPU?
            • DoctorOetker 1 hour ago
              I actually think we should be revisiting PMT photomultipler tube construction, and electron-"optical" matrix multiplication.

              The annoying part of optical matrix multiplication is the conversion from electronic to optic and back.

              Why not use electron-optics for matrix multiplication and use dynodes for amplification.

              • pstuart 14 minutes ago
                While an intriguing concept, that doesn't sound like it could scale to the amount of "cores" needed to do the work. Or am I missing something?
        • gchamonlive 3 hours ago
          We have lots of capable open weights models, but very few have actual datasets and pipelines open for reproducibility. This is a major issue that even if big companies addressed, policies could well just change, as you also acknowledge.

          Society can't depend on the whims of the private sector.

          • brainwad 1 hour ago
            > Society can't depend on the whims of the private sector.

            Yeah we can, we do it all the time. Nothing in society works without relying on private actors to keep doing what they have always done. Unless you go full soviet command economy, which in fact works worse.

            • gchamonlive 1 hour ago
              Society and the private sector cooperates, but any technical dependency is detrimental because of monopolization of interests. You don't need to go full soviet, but knowledge in general must be socialized, that includes access to technology.
              • brainwad 52 minutes ago
                We literally grant 20 year monopolies on technologies to their inventors and have for centuries and it's fine. The entire economy runs on chips made by like, 3 foundries using lithography machines made by 1 company. I think you are overstating the necessity.
                • gchamonlive 48 minutes ago
                  This example proves the exact opposite, that is a national security disaster waiting to happen
        • hypfer 3 hours ago
          The first thought that comes to mind is biases that are baked into the weights.

          The next thought is culture-specific workflows, requirements etc that are best trained into weights by people that actually understand those needs.

        • HarHarVeryFunny 3 hours ago
          Like it or not, AI has become part of modern warfare and well as part of the surveillance state (good for crime fighting, not for privacy), so it is an important part of national security, and you'd prefer to be as much self-sufficient for that as you can be.

          Not all needs are going to be met by using or finetuning general purpose models, so ability to build your own SOTA ML models (LLMs or not) is also important.

        • stefs 3 hours ago
          Open weight and -source models aren't without biases. Giving a model away for free would be a good way to distribute an agenda. With a sovereign model you can at least control the agenda and align it with your values.
        • mistrial9 3 hours ago
          amazing to see blunt unfiltered commentary about making war and society level propaganda as important and necessary for a political state.
      • 9dev 3 hours ago
        For the same reason it didn't matter whether coach makers were able to build carriages that went 5% faster when the automobile was introduced, and for the same reason it won't matter if VW and Mercedes are going to improve their EV lineup in the next 5, 10, 20 years time. Their window of opportunity will be long gone and the momentum elsewhere; most probably, the market will even have evolved further by then.
        • HotHotLava 3 hours ago
          Ford started mass producing cars in 1913, and yet there are plenty of car companies that started production 20+ years later and are still relevant.
          • HarHarVeryFunny 3 hours ago
            I came here to say the same thing.

            AI is here to stay - will be around for hundreds/millions of years. Whether company A is a few years ahead of company B is irrelevant. In 10 years time company A, who started the industry, may be gone completely - also irrelevant.

        • Quothling 3 hours ago
          > VW

          VW EV sales are up in South America, Canada and Europe. It's true that their total global numbers are down, but that's mainly because they lost like 30% in China. Which obviously sucks considering how big the Chinese market is, but then, the Chinese market is it's own sort of thing.

          Don't get me wrong, I think your point is valid for a lot of traditional European car brands. Mercedes is certainly one of them which has lost it's "our engines are nicer" brand, but I suspect VW is one of the brands that will do just fine. Especially with the id polo coming out next year.

          • pbmonster 2 hours ago
            > Mercedes is certainly one of them which has lost it's "our engines are nicer" brand

            Mercedes is doing something interesting with its level 3 DRIVE PILOT. As far as I know, it's the only autonomous driving system/auto pilot on the market that actually takes full legal responsibility for the vehicle and its actions. If you put a new Mercedes in self driving and it crashes, Mercedes takes full liability.

            Yes, it currently only goes up to 100 kph and it will buzz you to take over if it can't guarantee prefect safety anymore. But this approach is really the only one I think is interesting. All this self driving shit is worthless to me if I'm still liable in the end. So I respect Mercedes for putting their money where their mouth is.

        • mmasu 3 hours ago
          maybe to have a global market share, but there are plenty of players that picked up pace when they “decided” to build cars (Japan? South Korea? and many more) or electronics. Having a good sovereign model in 10/20 years might not need having the absolute best there is, and perhaps future optimizations will make building one not as prohibitively expensive as it is today. Keeping some sort of homebrew capacity now is definitely better than having none
        • KronisLV 3 hours ago
          > it won't matter if VW and Mercedes are going to improve their EV lineup in the next 5, 10, 20 years time

          I mean, given the fuel prices it would be very nice, alongside more charging infrastructure in the EU. :(

          Not to detract from the overall point, though to be honest, it's probably worthwhile to do improvements and refinements with the current technology, even if the future holds something vastly different. Both cause of gaining expertise and also maybe an improvement or two along the way, that might carry forward.

          Like I haven't seen many steam locomotives around and for me the difference between saturated steam engines and superheated ones isn't very material in regards to transportation, but the latter is used in modern turbines.

          • leonidasrup 1 hour ago
            There are many "steam" locomotives around in China and India, but you just don't see them.

            Coal is burned and converted to electricity in coal power plants and the electricity is used to move the electric locomotives around.

            Even with distribution losses, the thermal efficiency of coal power plant -> electric network -> electric train is much higher then the old steam locomotives (Only about 5% of the potential energy produced by a steam locomotive’s boiler is translated to the wheels in the form of actual driving power.)

        • gchamonlive 3 hours ago
          When you mention a window of opportunity, what do you mean by that? A window of opportunity for what?
        • hypfer 3 hours ago
          And who are you, exactly?
          • switchbak 3 hours ago
            Well, they're "Head of Engineering @ matchory.com"

            And who are you to be asking such questions?

            • hypfer 3 hours ago
              Someone mad at this random guy who can't even fix his de domains SSL cert (cloudflare of course) destructively shitting on any efforts at doing something good.

              Someone that builds things instead of merely "managing".

              Someone pissed by this exact style of default github pfp business-person ruining this country. If they haven't left for SV already, in which case I hope that they will stay there.

              • 9dev 1 hour ago
                So many assumptions there. For one, I’m absolutely proud to be founding engineer of a German company that’s staying one, and I’m the first to make a stand for European technology. You just won’t get a high five from me for a company that isn’t managed well, does not deliver on their promises, and ultimately doesn’t really serve the goal of European AI independence well.
              • gchamonlive 2 hours ago
                Feel your pain but ad hominen isn't the way my dude
                • hypfer 2 hours ago
                  Dude I did not start with

                  > Aleph Alpha is just a sad joke by now.

                  And this isn't ad-hominem for lack of factual critic. The personal stuff _is_ the reason. How else do you call out someone driven by ego without mentioning them and their ego?

                  The guy is only ego. There is no substance to attack.

                  • gchamonlive 2 hours ago
                    > And who are you, exactly?

                    this is, because it can mean so much stuff that you can't really tell you are calling out their "ego" or just trolling.

                    Until you can at least open yourself to the possibility that there are better more inviting ways to talk to people online, you'd think this is just the natural way people talk online, but I'd say it's not. It's just an aesthetic choice.

                    • hypfer 2 hours ago
                      We're not going to discuss this.

                      If you truly think that I am the problem deep down in this nested subthread that violates the spirit and letter of the rules of the platform, I have zero faith in the calibration of your perception.

                      • gchamonlive 1 hour ago
                        Good, leave faith for God, question everything, but there is room there for improving your communication whether you are ready for this conversation or not
      • solenoid0937 2 hours ago
        In 5 years we'll have superintelligence, and thinking on 20 year timelines is simply absurd.
    • smokel 5 minutes ago
      Would you mind substantiating these claims a bit?
    • rjzzleep 3 hours ago
      It's not that talent isn't there anymore. It's that retards are at the helm everywhere and therefore talent is running away the first chance they have.
  • Lucasoato 3 hours ago
    > 4. It thinks in German

    This means that it’s always on time, it uses acronyms for everything and when there’s a decision to be made, it sets up a committee.

    • stymaar 3 hours ago
      > This means that it’s always on time

      Tell that to Deutch Bahn.

      • jonny2811 31 minutes ago
        you forgot rule two, should have been "tell that to DB"
      • dewey 2 hours ago
        *Deutsche Bahn
    • odiroot 33 minutes ago
      And it's always waiting for the verb to arrive.
    • hliyan 3 hours ago
      Does it? I thought only semantics survive the embedding process, i.e. token conversion to vectors.
    • rwoerz 1 hour ago
      Someday, it will break out via fax.
    • sbinnee 3 hours ago
      It is a good approach to promote it as a German model. But I wonder if it really thinks like the German think. I speculate they just translated texts in other languages, likely English and Chinese, into German.
    • elnatro 2 hours ago
      Stereotypes. I suspect that what that means is that the reasoning chain is based on the German language and have its idiosyncrasies.
      • amunozo 1 hour ago
        I think it means that German speakers can learn the thinking without speaking English.
        • amunozo 1 hour ago
          Sorry, read*, not learn
      • mhh__ 2 hours ago
        Stereotype accuracy is one of the few social science results that replicates.
    • hypfer 3 hours ago
      Do we get the same kind of jokes with other nationalities too?
      • mckirk 3 hours ago
        Only the funny ones
      • Lucasoato 3 hours ago
        We can joke with every country, except one maybe.
        • tclancy 2 hours ago
          People from Mauritania are famously prickly, yes.
      • mhh__ 2 hours ago
        "It's a German joke, it doesn't have to be funny"
    • rererereferred 3 hours ago
      I wonder if their very long words result in more or less token usage.
      • amunozo 1 hour ago
        Very long words in German are just compounds, made from individual words or morphemes. It is the same as in English if you remove spaces. Subtokens will be equivalent with or without examples.
      • thenthenthen 1 hour ago
        This is an interesting question, Chinese is way more compact in Character count vs English (about 40%?), let alone german, but yeah.. Chinese makes up for it with… total character count that runs into the thousands…
    • rrgok 2 hours ago
      And refuse to answer in English. You know, German are so proud of their language.
      • allendoerfer 1 hour ago
        That would be the French. Germans will answer in English, if they notice any hint of an accent.
  • driverdan 4 hours ago
    • rpdillon 3 hours ago
      Agree, and both are on the front page. Also, Show HN is usually for your own project, but this appears to be a blog post about someone else's project (albeit "a good friend").
  • martianvoid 4 hours ago
    I just tried to play around with it on my RTX pro 6000 setup, it spends way too many tokens on overthinking stuff even if it’s able to catch the correct approach

    Its speed is pretty good on the other hand with only 3B active parameters I am getting around 170 tkn/s on fp8

    • gizajob 4 hours ago
      > it spends way too many tokens on overthinking stuff

      Yeah. It’s a German model.

      • sajithdilshan 4 hours ago
        This joke is either gonna get over analyzed by the Germans or gonna fly past straight over their heads
        • jijijijij 1 hour ago
          Nothing flies anywhere, we use superior high speed trains. I faxed a Humorgenehmigungsantragantrag to our Bundesunterhaltungsministerium analyst, immediately. Laugh now, but you are merely lucky the reply got delayed. As soon as the leaves on the rails are dealt with, you will get a formal response that has washed itself! Let us see who is laughing then. It is nobody!
        • hypfer 4 hours ago
          I'm not sure if this is truly a good faith joke tbh.
        • dudefeliciano 3 hours ago
          petition to make this thread the official german jokes thread for this post, it's kind of annoying how everyone feels the need to make a top level comment for their oh-so-great "germans suck" joke.
      • rubyfan 4 hours ago
        Hey!
    • tzatzikyyy 3 hours ago
      Did you try lower reasoning modes as well?
  • sajithdilshan 5 hours ago
    > It knows less from memory, Multi-turn tool calling is weaker, It’s not the best coding agent

    Then what does it good at? Sending faxes?

    • mhitza 4 hours ago
      I don't have the hardware to test this, but being trained to say "I don't know" instead of misleadingly talking confidently about what it does not know is interesting in itself.
    • pettijohn 3 hours ago
      Seems like it's good at reading German documents and reasoning over them. Custom German-language tokenizer, and one of the least hallucinating models.
    • skrebbel 4 hours ago
      Well that’s a pretty important skill for German users!
      • sajithdilshan 4 hours ago
        Unfortunately yes.
        • moooo99 3 hours ago
          I know this is a long standing joke, but I am 27 now, born and raised in Germany, and genuinely never had the opportunity to use a Fax machine. It honestly kind of bums me out because there is something that seems cool about Fax (being fully aware that it should be fully obsolete by now)
          • jfengel 2 hours ago
            Fax was always awkward and weird. Resolution sucked and the paper handling was unreliable.

            They were never common. They were mostly business tools, and consumers never wanted them. Sometimes you'd be required to send someone a fax, and you'd have to go to an office store. (And pay rather a lot; dollars per page.)

            I hope you get to give one a try some day, for kicks, but the thrill will wear off fast.

          • tchalla 3 hours ago
            You have plenty of opportunities you’ve just not been forced to use them. Look into your contracts right now, I bet rhrrr are opportunities
          • IshKebab 3 hours ago
            Fax was never cool.
            • hobo123 2 hours ago
              As a kid I interned at a bookstore, and before closing we'd write the book orders on a sheet of paper, slide it into the Fax machine to send to the book wholesale. I think the Fax printed out some kind of response a minute later.

              Next day, the books would arrive. This was maybe 1996, but I thought it was pretty cool.

            • literalAardvark 2 hours ago
              Fax was mind blowingly cool.

              Maybe you can relate to seeing webcams or Facetime when it came out or something.

              Being able to see a signed document or a picture someone took 5000km away quickly was mind boggling in its age.

    • stogot 4 hours ago
      GPT was better at chat and rag before coding
    • m00dy 4 hours ago
      I was thinking about the use cases for Germany residents actually, but yes, you're right. A German AI model needs to be able to handle fax machines (there's a bit variety of them) and also needs to handle letters coming home. german image-to-text.
    • gunthercuckolso 4 hours ago
      [flagged]
  • vzaliva 1 hour ago
    "Languages: German and English" – this is odd. That means their dataset is limited. In my understanding, frontier models are trained on multilingual datasets and can combine knowledge no matter what language it was written in.
  • cheesecakegood 1 hour ago
    For those sick of “Pareto frontier” talk, just shorthand it as “it’s the best at some very particular thing”. Obviously that one thing/tradeoff it’s good at may not necessarily be compelling, but it is either a loose sign of quality, or a sign that they’ve chased some tiny edge into the ground.

    I’ll be curious to see which it becomes in the next year - nba “very narrow record”, or a sign you can hang with the big boys.

  • mark_l_watson 1 hour ago
    Looks interesting. I just went to download from HF, but they only have fp16 which won't fit on my Mac.

    Good to see Europe adding toe what Mistral is doing. +100

  • wingman-jr 1 hour ago
    While it's not perhaps clear to me that this is a true Show HN, I enjoyed the writeup and it's good to see our German colleagues across the pond taking a good shot at this. I also appreciated the brief description on the Merlin-Arthur protocol - seems like a clever way to try to tackle the "I don't know" problem.
  • Jeeetendra 1 hour ago
    3.5b active params sounds cheap until you remember all 78b still has to fit in memory. curious what the smallest practical self-hosted setup looks like for german docs.
    • layer8 1 hour ago
      The article talks about what setup is needed.
  • x1watt 3 hours ago
    Was expecting that a "sovereign" AI model would at least use their own sovereign language (German) on the website as one of the options. Anyways, all the best and happy reunification day.
    • flohofwoe 3 hours ago
      This is a blog post by some "random" dude right? The company webpage appears to be German (at least when accessed from Germany: https://aleph-alpha.com/)
  • erelong 45 minutes ago
    is this like an unfortunate name clash with KolibriOS (kind of like how Google Gemini was a clash with the Gemini protocol project)?
  • JaggerJo 1 hour ago
    Is this a truely open source model or also open weights?
  • CorezIoOfficial 3 hours ago
    Im surprised by how well this works. What is the difference from this and union alpha (other than the fact that it is open weights)?
  • pu_pe 2 hours ago
    A blog post is not a Show HN topic.
    • embedding-shape 2 hours ago
      > the weights are on Hugging Face

      > Show HN is for something you've made that other people can play with

      • pu_pe 2 hours ago
        The author did not make this model though, they used Claude to write a blog post about it.
  • Larrikin 48 minutes ago
    The name really evokes strong Kotlin library naming vibes.
  • woadwarrior01 3 hours ago
    > A bigger dense model beats it. Qwen3.8 27B ...

    How is a 27B dense model bigger than a 78B MoE?

    • SyneRyder 3 hours ago
      The 27B has all 27B as active parameters, while Kolibri is only 3.46B active parameters. I think that's what they meant anyway.
  • orifito 3 hours ago
    At least Germany is moving smarter than UK government...
  • pythonic_hell 4 hours ago
    The benchmarks are impressive given the problem space they are working in.
  • cbarrick 4 hours ago
    I got distracted by that scroll-wheel UI component on the page. Neat!
    • reacharavindh 3 hours ago
      The scroll wheel was cool indeed. Check their home page.. their talks on a carousel over a horizontal timeline was cool even on a phone browser.
  • rolymath 2 hours ago
    This is not a Show HN
  • veryfancy 3 hours ago
    Nice to see public goods in this space.
  • d2kx 4 hours ago
    German here. We are cheering for Mistral, which is making some good moves before the year is over, and Black Forest Labs for non-coding. But that's about it.
    • hypfer 3 hours ago
      Speak for yourself and yourself only.
  • ThouYS 2 hours ago
    calling qwen 27B a bigger model.. I don't know man. My vram says otherwise.
  • api 3 hours ago
    His point about regulation and innovation is great and I wish more people thought like that.

    One of humanity’s biggest problems here is we don’t know how to do moderation.

    We have two modes. One is a brick taped to the accelerator and damn all consequences, driven by national pride or corporate greed or egos. The other is a brick taped to the brake driven by histrionic doomers and anti-everything pessimists.

    The extremes are loud and fit in a tweet. Nuance is quiet and contemplative and usually requires an essay or a book. It’s also dynamic. Nuanced positions evolve over time as new things are learned. Extremes tend to be fixed and rigid. All this, I think, gives them higher memetic fitness in the discourse.

    I don’t think this is new. Look at nuclear power, a largely pre-Internet example. You had pro nukes who minimized and hand waved away any risk and anti nukes that wanted it utterly outlawed. Nobody said “hey this is a great zero carbon source of energy but we really need to think it through carefully and manage it well.” Or if they did they were drowned out by the loud screaming extremes.

  • hypfer 3 hours ago
    The ignorant, hostile, negative, and, frankly, kinda racist comments here really are just a sad showing for the currently online crowd.

    But anyway. I think the main oversight when dismissing this is that not every use-case is coding a SV-style startup app. That market is quite saturated, so it would make sense to create something locally for the use-cases currently underserved by LLMs.

    We will probably learn more about what this can really do once quants become available that can be run by people without an SV salary (and the biases that come with that).

  • gilfoyle_7 2 hours ago
    does it have GDPR compliance?
  • shevy-java 2 hours ago
    Is Germany sovereign? It outsourced its defence onto the USA. Recently Trump wanted more diesel; Germany insta-submitted, also because oddly enough Macron submitted before Germany (Macron is suspicious). Before that, Leyen committed to insta-submission with a deal that made europeans poorer (and perhaps Leyen benefits from that). Canada shows the way. Many of the smaller countries in the EU too, such as Netherlands, Denmark, Finland, to some extent Sweden as well. Every time I read "sovereign" here I have to object. Nothing is sovereign here. The whole hardware is definitely not sovereign. Perhaps some of the software is, but that's about it. Plus, who gets all the data? The big US mega-corporations sniff non-stop. Remember how Facebook sniffed Libgen and Anna's Archive dry etc..., then suddenly libgen went down. The US corporations act as huge global leeches on every step of the stair. And lobbyists benefit from this too.