There’s been this weird idea lately, even among people who used to recognize that copyright only empowers the largest gatekeepers, that in the AI world we have to magically flip the script on copyr…
How do you know what it’s storing? I certainly don’t, but I know what the security researchers have found that proved it was storing copyrighted material and real people’s private info or PII.
You being able to spit people’s name and personal details doesn’t mean you are keeping a database of those details in your brain. It’s all just neurons and the connection between them that can be triggered to extract those details out.
LLMs also attempt to mimic this method of not storing direct information, but tweaking parameters to ‘learn’ the information. Inside LLMs are just a bunch of parameters that if not well-designed, can be made to spit out what they have learnt. That doesn’t mean they store those information as is.
It’s not just what they tell you. There are plenty of publicly accessible LLM models. Go and download them and open the files up. Surely if they are storing these things as complete data, you can easily find them by poking around the files instead of having to make then spit it out.
I’m aware of the availability of them, I’ve looked into building a private install of GPT4All. Even though we can look into those files directly, it doesn’t prove that the large “AI” systems run by the mega-corps are not storing copyrighted data. The only thing that could prove that is a complete audit of all the data storage that their “AI” systems have access to.
This will likely play out in the courts due to the numerous lawsuits in process from artists suing over their work being stolen. Legal discovery could compel that kind of data audit.
How do you know what it’s storing? I certainly don’t, but I know what the security researchers have found that proved it was storing copyrighted material and real people’s private info or PII.
You being able to spit people’s name and personal details doesn’t mean you are keeping a database of those details in your brain. It’s all just neurons and the connection between them that can be triggered to extract those details out.
LLMs also attempt to mimic this method of not storing direct information, but tweaking parameters to ‘learn’ the information. Inside LLMs are just a bunch of parameters that if not well-designed, can be made to spit out what they have learnt. That doesn’t mean they store those information as is.
That’s what they tell you about it I’m sure, but what proof do you have?
It’s not just what they tell you. There are plenty of publicly accessible LLM models. Go and download them and open the files up. Surely if they are storing these things as complete data, you can easily find them by poking around the files instead of having to make then spit it out.
I’m aware of the availability of them, I’ve looked into building a private install of GPT4All. Even though we can look into those files directly, it doesn’t prove that the large “AI” systems run by the mega-corps are not storing copyrighted data. The only thing that could prove that is a complete audit of all the data storage that their “AI” systems have access to.
This will likely play out in the courts due to the numerous lawsuits in process from artists suing over their work being stolen. Legal discovery could compel that kind of data audit.