I think it would be totally incoherent to say that copyrighted information is essentially fair game to include in your model but the outputs of your model are privileged against being included in other models.
I'm not sure "competition" is necessarily the right word but I don't think that model creators and their political backers have a leg to stand on when complaining about distillation.
One is allegedly a copyright violation (fair use IMHO but still a grey area legally AIUI), one is a very explicitly a violation of the terms of service you agree to when signing up. And yes I do think that if somebody puts their website behind some barrier where a user has to actively affirm they won't download the contents and use it to train an LLM that should be protected by the same law for the same reason.
Are you arguing that training on copyrighted material is fine but distilling against tos is somehow bad? Lmao what a take. Nah, if copyright cant even be enforced then theres no way on earth violating a tos should ever have repercussions other than they ban you. stop carrying water theyre not gonna pay you
It's so nice of you to think of the poor ignored Terms of Use rather than merely clicking the nag box to proceed like everyone else. Maybe with some more positive attention the Terms might become less depressed, abstain from binge eating, and stop putting on pounds of new paragraphs every year.
Exactly. AI companies distilled the whole internet into AI weights, now other AI companies are distilling their work to generate new AI weights. Nothing wrong with that, and the world gets open weights AI models.
> I think it would be totally incoherent to say that copyrighted information is essentially fair game to include in your model but the outputs of your model are privileged against being included in other models.
IANAL, but going from copyrighted texts to an AI model likely constitutes a new original creation, not a slavish reproduction. Whereas going from one AI model to another is more likely to be considered a slavish reproduction reproduction, although arguably AI models are not copyrightable to the extent that the weights are objective facts, like entries in a phone book.
Downloading and redistributing companies' exact AI models would violate copyright. Querying multiple AI models to form your own model with a unique set of weights does not.
Realistically the companies will organise their data through a network which will involve distributing and storing the copyrighted material without the consent of the copyright holder. This is in and of itself generally ruled to be a breach of copyright irrespective of whether it is fed into a model or not.
Saying AI output is public domain is a misinterpretation of that case. The only court case related to the monkey selfies was by PETA, and the court ruled that the monkeys don't hold the copyright. The court didn't rule on who, if anyone, holds the copyright; whether Slater holds the copyright or it's public domain is undecided. The official opinion of the US Copyright Office is that generative AI outputs can be copyrighted by a person "to the extent that their contributions qualify as authorship".
Yes, copyright requires a human - in much the same way that an LLM requires a prompt. There's a continuum of author involvement that has to be evaluated on a case-by-case basis.
It's worth noting in your case that no determination of copyright was actually made by a court (well, almost: a court ruled that the monkey definitely didn't have copyright and that even if it did they wouldn't make PETA its guardian).
Some opinions were issued but those are non-binding and eventually the photographer gave up because litigating copyright is expensive.
I'm not sure I can defend my comment. So I'm feeling like taking it back.
Regarding CEOs - it looks pretty common yes.
I think rule breaking is a quality I've seen in entrepreneurs, thinking differently to use that euphemism. Challenging the dominant paradigm, to be tongue-in-cheek.
I imagine it's a quality engendered by business leadership schools or maybe Market economics? To Find a need and fill it, but taken to the extreme. Can be figuring out where the illegal line is and walking it.
Next, Commenting on the actual article again ..
China, the Chinese government has a known history of stealing technology from other nations. (Aside: I assume all governments do this, State-Sponsored espionage/industrial or whatever.) So that is actual theft. China is also supporting industries that it considers strategically important to success. I am ignorant at the difference between, or limitation of reach of a Chinese government supported influence of theft activities, versus a business that happens to be in China and exhibits theft behavior All on its own.
The area that I'm puzzling about is noticing how this business leader is in a position, through extraordinary concentration of wealth and power, to influence the efficacy of state-sponsored industrial espionage by helping to propose state-level policies by the United States and internationally. I don't know if the word oligarchy is the right word, but, it's like a corporation can author international financial and political policy. Pretty wild considering Nvidia was just like a $80 stock back in in the late '90s.
I have no comment one way or the other, other than to say that I really respect this kind of attitude in discourse. Realizing you can’t defend an argument, and choosing not to just dig in and fight, is exactly the kind of thing that makes us better as a society. I hope I can be more like you!
I agree that AI companies shouldn’t be allowed to bar distillation under grounds of fairness and societal benefit.
But I don’t see how copyright has anything to do with anything. Copyright is a legal concept, not an ethical concept. Training an LLM on copyrighted material is legal. Ethically, whether a work that the LLM trained on is copyrighted or not has no relevance
The accusation of theft by anyone is really only as strong as their financial capability to enforce it. The leg that the model companies stand on is their bank account, which is increasingly becoming the only force that moves the needle in this economic and political environment.
If your recommendation algorithm always recommends B when it sees A, and I derive this information from exfiltrating your masked data , is it not theft simply because I didn’t unmask the true identity of A and B?
Stole all other important information embedded in the relationship (the essence of the relationship itself).
This is not seen as theft to only two or three types of people:
1) Technically ignorant
2) Or Technically ignorant and morally bankrupt
3) Truly evil, a combination of technically capable, and morally bankrupt.
(An “Abomination” is also possible, unaware and unable to perceive morality , not just mere ignorance. For example, it is tremendously generous to call an Abomination ignorant or amoral, that’s a compliment to such a thing. Would Jensen Huang know about the words coming out of his mouth, is the true question really.)
I think that Jensen's position as the guy selling the proverbial shovels incentivizes him to take a lot of irresponsible positions (namely around safety) but imo he's fundamentally correct here.
>he doesn't think kids should learn their multiplication tables
He said „The majority of the kids shouldn't learn”, not everybody. I listened to that pod and was under the impression that he wants to be portrayed as a "grounded" person and i didn't find his takes to be out of touch.
He comes off bad in this interview. Out of touch is correct - "I had to pump gas once a few years ago ..." and panicked because he didn't know his address or zip code.
Personally I think multiplication tables aren’t that useful because they’re just rote memorization. I think it’s better to teach kids to memorize a few such as x * 2, x * 5, x * 10 and teach them how to extrapolate from there. Extrapolation is a more useful skill than memorization.
I don’t think the rote memorization step per se is important, but the knowledge is. You can pick it up just from doing a lot of math work with those basics. Same thing in chemistry - you will learn most of the periodic table automatically just from doing so many lookups.
But IME, not having the knowledge is a major handicap. I did eventually learn essentially the full multiplication table with speedy recall and it _is_ useful in day-to-day life. I think its basic numeracy on a par with literacy if you ever have to do basics like compare prices for consumer goods, sanity check claimed facts/statistics, etc. I don’t think those skills go away when you have AI anymore than they do when you can “trust” Amazon to do the price math for you.
The tables have normally only gone to {12}x{12} and most young children have no issue. Why would we cut back on this?
The memorization means that you can extrapolate the more complex concepts. If you don't have the very basics, what are you even going to extrapolate on?
How much of your preferred programming language have you memorized?
How has this enabled you to engage in higher-level thinking instead of syntax correctness?
(I saw this pattern first hand with 100s of students while teaching intro cs)
---
my current favorite definition for The Singularity is the point at which humans give up critical thinking and have decided deferring to Ai is the better idea, social media being a stepping stone on that path that primed us for the outsourcing
many of the commentary under this submission has me thinking we are well on our way in this journey
As sometime who's memorized but also forgotten their timetables while still managing to go through Algebra I. I don't understand the importance for memorizing times tables (or any other tables for that matter).
Most people on a budget (i.e. the vast majority of people) are constantly using exact or approximate multiplication while shopping. And they are rarely whipping out their calculator for this purpose.
Why not just use the calculator app on your smart phone? That's what I do. (School made me all too aware of how easy it is for me to get the wrong answer when trying to do calculations mentally)
Depends on the calculation and context, it is definitely advisable to use tools for complex or high-risk situations, dividing the dinner bill by 2 or 3, that should not require a calculator or one of the thousands of bill splitting apps, perhaps it is a bit harder when you want divide bill by items, but not a high risk situation.
practice makes perfect, exercise your brain like your body
There are basic skills that everyone should learn, being able to do simple math in your head is a building block for more complex reasoning.
I once had a roommate who I encouraged to divide our grocery bill by 2 in his head for splitting the bill, by the end of the year he was quick and accurate.
That doesn't follow. I'm not especially gifted at mental math yet I've been able to learn various programming languages and libraries and dev tools. I've also been able understand and implement some algorithms.
Quick math. Basic financial literacy. Times tables are for arithmetic, the working man's math. Algebra is an entirely different math which solves different problems and deals more in the abstract.
You can definitely get past algebra without knowing your times tables, but algebra isn't going to help you in the checkout line to make sure you received the correct amount of change and bought the correct number of goods.
AFAIK he was outright lying when he said he forgot his zip code. It hasn’t been possible to buy gas in Bay Area for decades without inputting your zip code for credit card purchases. That smelled of some bad speechwriting.
Now when the big LLM companies create their products by accumulating a notable portion of copyrighted creative works (including computer programs) from Internet, it does not count as copyright infringement or competition (or "theft"). It is considered as "fair use". So, why training LLMs on other LLMs is not fair use too?
> If you don’t like that, if you don’t like people to use your products, all you [have to do is] know your customers, and disable the service
Considering the current situation, I understand Mr. Huang point, but that's not how the copyright framework assumed to be working from the beginning. It should protect both small actors (authors) and the big companies from unrestricted use of creative works. Now this mechanism seems to be practically dysfunctional.
And a big portion of this lies on shoulders of proponents of permissive OSS, who defend an idea of (almost) unrestricted use of their source code texts for many years, and long before mass LLM scrapping became a thing.
I wonder how the "Chinese AI is all just distillation" folk are going to cope now that the Chinese are all aboard the post-train via agents in custom training environments wagon?
Here's Xiaomi's discussion of this, plus their open-sourcing of 7000+ RL training environments.
Can someone answer this: if anthropic cares so much, why don't they meganerf their thinking summaries (edit: in Claude Code) the way OpenAI does? OpenAI's thinking traces so much that the Codex app doesnt even show them.
In Claude Code (at least inside VSCode) thinking traces are hidden by default. You can reenable them via JSON. They’re not the real thinking traces though, they are something like thinking summaries.
It's not there in the chatui but claude code shows more reasoning than they need to.
im not suggesting that they should, but it's so much more than the garbage both antigravity and codex show to users. seems like a thing a company paranoid about distillation would reign in.
'Fair play' is a dimension of competition. They aren't orthogonal attributes.
You can compete fairly or unfairly.
In any event, distillation seems to be in a grey enough area that reasonable people can disagree (IMO). It strikes me as more akin to theft than to fair competition, but it doesn't seem like it should be treated as criminal. Rather, I think the onus should be on the labs to defend themselves against it.
All that said, Jensen's take is clearly quite biased.
Hasn’t it already been ruled by a court that LLM outputs are not subject to copyright? It opens a can of worms to say that LLM companies can dictate what their outputs are used for.
I think copyright is mostly not relevant. Contracts operate mostly independent of copyright law (in the U.S.). If OpenAI puts in their license something to the effect of "you may not train your LLMs on these outputs" and/or "your access is limited in these ways", but you violate the license, they can and should block your access and sue you. These are things that can be monitored and enforced under existing law.
And, in fact, they already do this:
- OpenAI [1]: "[you may not] Use Output to develop models that compete with OpenAI."
- Google [2]: "You may not use the Services to develop machine learning models or related technology."
- Anthropic [3]: "[You may not use our services] to develop any products or services that compete with our Services, including to develop or train any artificial intelligence or machine learning algorithms or models or resell the Services"
Jensen is borderline comical on his instances of high GPU usage. OpenClaw? "let's do it, it's amazing!", selling GPUs to china "it's of uttermost importance, USA can't afford not selling GPUs to china because other will do", chinese firms burning GPU cycles like hell on american datacenters from openai and anthropic? "this is just competition, let's do this, this is amazing".
On paper, but the world isn’t constrained by such trivial matters when enough power is behind you. Although, smuggling only goes so far, obviously, it’s not like they’re not available.
No. If a CEO makes self-serving comments that are in their interests, but not the company’s interests, that could be fiduciary misconduct. With that being said, CEOs have a fair amount of latitude in what they can say, since it’s difficult to tell what the “optimal” thing to say at any point would be (see: Tim Cook telling an activist investor to “get out of the stock” if he didn’t like Apple’s environmental initiatives).
The idea that CEOs’ hands are tied, and that they are kept on some sort of short legal leash, is largely (though not entirely) a fiction.
I'm not sure "competition" is necessarily the right word but I don't think that model creators and their political backers have a leg to stand on when complaining about distillation.
IANAL, but going from copyrighted texts to an AI model likely constitutes a new original creation, not a slavish reproduction. Whereas going from one AI model to another is more likely to be considered a slavish reproduction reproduction, although arguably AI models are not copyrightable to the extent that the weights are objective facts, like entries in a phone book.
It's entirely possible for AI output to be copyrighted if it meets the requirements; prompting an AI takes some skill.
https://en.wikipedia.org/wiki/Monkey_selfie_copyright_disput...
Acquiring and setting up a camera for the right shot takes some skill. That does not matter.
It's worth noting in your case that no determination of copyright was actually made by a court (well, almost: a court ruled that the monkey definitely didn't have copyright and that even if it did they wouldn't make PETA its guardian).
Some opinions were issued but those are non-binding and eventually the photographer gave up because litigating copyright is expensive.
Original:
Smells like a CEO justifying theft. It is reasonable to believe Wong thinks differently about law breaking, given this article.
Something something legal liabilities something something?
Regarding CEOs - it looks pretty common yes.
I think rule breaking is a quality I've seen in entrepreneurs, thinking differently to use that euphemism. Challenging the dominant paradigm, to be tongue-in-cheek.
I imagine it's a quality engendered by business leadership schools or maybe Market economics? To Find a need and fill it, but taken to the extreme. Can be figuring out where the illegal line is and walking it.
Next, Commenting on the actual article again ..
China, the Chinese government has a known history of stealing technology from other nations. (Aside: I assume all governments do this, State-Sponsored espionage/industrial or whatever.) So that is actual theft. China is also supporting industries that it considers strategically important to success. I am ignorant at the difference between, or limitation of reach of a Chinese government supported influence of theft activities, versus a business that happens to be in China and exhibits theft behavior All on its own.
The area that I'm puzzling about is noticing how this business leader is in a position, through extraordinary concentration of wealth and power, to influence the efficacy of state-sponsored industrial espionage by helping to propose state-level policies by the United States and internationally. I don't know if the word oligarchy is the right word, but, it's like a corporation can author international financial and political policy. Pretty wild considering Nvidia was just like a $80 stock back in in the late '90s.
But I don’t see how copyright has anything to do with anything. Copyright is a legal concept, not an ethical concept. Training an LLM on copyrighted material is legal. Ethically, whether a work that the LLM trained on is copyrighted or not has no relevance
Stole all other important information embedded in the relationship (the essence of the relationship itself).
This is not seen as theft to only two or three types of people:
1) Technically ignorant
2) Or Technically ignorant and morally bankrupt
3) Truly evil, a combination of technically capable, and morally bankrupt.
(An “Abomination” is also possible, unaware and unable to perceive morality , not just mere ignorance. For example, it is tremendously generous to call an Abomination ignorant or amoral, that’s a compliment to such a thing. Would Jensen Huang know about the words coming out of his mouth, is the true question really.)
while I generally agree on his open weight stances, I lost all respect in that moment, everyone of these people are so out of touch
He said „The majority of the kids shouldn't learn”, not everybody. I listened to that pod and was under the impression that he wants to be portrayed as a "grounded" person and i didn't find his takes to be out of touch.
But IME, not having the knowledge is a major handicap. I did eventually learn essentially the full multiplication table with speedy recall and it _is_ useful in day-to-day life. I think its basic numeracy on a par with literacy if you ever have to do basics like compare prices for consumer goods, sanity check claimed facts/statistics, etc. I don’t think those skills go away when you have AI anymore than they do when you can “trust” Amazon to do the price math for you.
The memorization means that you can extrapolate the more complex concepts. If you don't have the very basics, what are you even going to extrapolate on?
It does not mean that. It means you can parrot a number if it came from the memorized set.
How has this enabled you to engage in higher-level thinking instead of syntax correctness?
(I saw this pattern first hand with 100s of students while teaching intro cs)
---
my current favorite definition for The Singularity is the point at which humans give up critical thinking and have decided deferring to Ai is the better idea, social media being a stepping stone on that path that primed us for the outsourcing
many of the commentary under this submission has me thinking we are well on our way in this journey
practice makes perfect, exercise your brain like your body
I once had a roommate who I encouraged to divide our grocery bill by 2 in his head for splitting the bill, by the end of the year he was quick and accurate.
You can definitely get past algebra without knowing your times tables, but algebra isn't going to help you in the checkout line to make sure you received the correct amount of change and bought the correct number of goods.
> If you don’t like that, if you don’t like people to use your products, all you [have to do is] know your customers, and disable the service
Considering the current situation, I understand Mr. Huang point, but that's not how the copyright framework assumed to be working from the beginning. It should protect both small actors (authors) and the big companies from unrestricted use of creative works. Now this mechanism seems to be practically dysfunctional.
And a big portion of this lies on shoulders of proponents of permissive OSS, who defend an idea of (almost) unrestricted use of their source code texts for many years, and long before mass LLM scrapping became a thing.
Here's Xiaomi's discussion of this, plus their open-sourcing of 7000+ RL training environments.
https://www.alphaxiv.org/abs/2609.mimo-scaling-reinforcement...
https://mimo.mi.com/docs/en-US/news/latest/v2-6
Who needs a few of someone else's "vacation postcards" of their post-training experience when your agents can go on vacation themselves!
im not suggesting that they should, but it's so much more than the garbage both antigravity and codex show to users. seems like a thing a company paranoid about distillation would reign in.
You can compete fairly or unfairly.
In any event, distillation seems to be in a grey enough area that reasonable people can disagree (IMO). It strikes me as more akin to theft than to fair competition, but it doesn't seem like it should be treated as criminal. Rather, I think the onus should be on the labs to defend themselves against it.
All that said, Jensen's take is clearly quite biased.
And, in fact, they already do this:
- OpenAI [1]: "[you may not] Use Output to develop models that compete with OpenAI."
- Google [2]: "You may not use the Services to develop machine learning models or related technology."
- Anthropic [3]: "[You may not use our services] to develop any products or services that compete with our Services, including to develop or train any artificial intelligence or machine learning algorithms or models or resell the Services"
[1] https://openai.com/policies/row-terms-of-use/
[2] https://policies.google.com/terms/generative-ai/archive/2023...
[3] https://www.anthropic.com/legal/consumer-terms
https://www.reuters.com/world/china/chinas-deepseek-trained-...
Why should Ai be different?
Fortunately open weights are likely to dominate and it becomes a permissionless ecosystem
The idea that CEOs’ hands are tied, and that they are kept on some sort of short legal leash, is largely (though not entirely) a fiction.
His company has been public longer than you’ve been solvent.
The shareholder part is also not correct in they way it's being implied.