This is like saying that smoking weed is unethical because of the murderous actions of the cartels who grow and distribute it. You’re problem with AI isn’t with the technology itself nor its use.
You’re problem with AI isn’t with the technology itself nor its use.
I’m not sure about the weed argument, can’t say anything about it. But first, my personal problem with Ai doesn’t matter to this disussion, as I was talking about what people could have an issue with, not about my personal problems. Secondly farming data without respecting the original license, not even linking to the source where it got from is in fact a problem with “its use”. That is an unsolved problem lot of people just ignore, which does not make it less of a problem (it makes it a bigger problem). And that is just a few issues I listed, not even talking about the problems of generated content. There are legitimate concerns using Ai, no matter how you frame it, some just choose to ignore it.
But first, my personal problem with Ai doesn’t matter to this disussion, as I was talking about what people could have an issue with, not about my personal problems.
Okay. Then the problem that people, not you particularly, have with AI isn’t with the technology itself nor its use, but the methods by which the companies “grow and harvest it”, so to speak.
Which I don’t even think is true. People have just lost their minds over this. If a company unethically steals data and publishes a book with that data, we don’t say that there are legitimate concerns using books.
We would be having a problem with it if every book was regurgitated stolen data. That’s kind of the core of how all of these generative AI models work to the point that their creators have argued in court that they wouldn’t exist at all if they weren’t allowed to steal all the training data.
It can be trained on the exact same free data that we can be trained on. Just like us, it doesn’t have to bypass paywalls. I think they’re arguing that if I can look at the Wikipedia and tell you that an avalanche in Pakistan killed ten climbers, then AI should be allowed to as well.
I think they’re arguing that if I can look at the Wikipedia and tell you that an avalanche in Pakistan killed ten climbers, then AI should be allowed to as well.
This is misleading and dishonest. Its not just looking up a single free information. Ai does more than that. Ai will not only scrape and copy the entire article, it will also scrape and copy the entirety of Wikipedia. There are licenses attacked to the Wikipedia article we have to follow to, if we scrape and use it for other purposes. It will also scrape the entire web. That is not the same as just looking up an Avalanche in Pakistan killed ten climbers.
This is weirdly accusatory. I’m not sitting here in a black cape and stovepipe top hat, twirling my mustache.
Its not just looking up a single free information. Ai does more than that. Ai will not only scrape and copy the entire article, it will also scrape and copy the entirety of Wikipedia. There are licenses attacked to the Wikipedia article we have to follow to, if we scrape and use it for other purposes. It will also scrape the entire web. That is not the same as just looking up an Avalanche in Pakistan killed ten climbers.
If this is the way we’re going to look at it then I guess we have to throw this technology out the window. It seems like a bad idea to do that though.
BTW so you agree with me then, because the only solution you see here to throw it away. Because what I described is not hallucinated by me, this is how its operate. It was to correct you for your wrong example with looking up a single information from Wikipedia.
we have to throw this technology out the window. It seems like a bad idea to do that though.
Why? Besides that’s not possible in the near future, it could be made illegal and then suddenly the biggest companies (the driver behind it) would have to stop. This wouldn’t eliminate all Ai tools everywhere, so getting rid of it is probably not possible anymore.
But that wasn’t my suggestion either. And we don’t have to throw it away, we have to comply. It needs to change.
I think you need to read up on the difference between “free as in beer” and “free” licenses. Everything on Wikipedia is copyrighted under a GNU license, which requires any use to give credit AND ALSO be licensed under the GPL. Which AI models do not (and cannot) do.
Also, just conceptually… We built a giant, free library (and museum and art gallery and etc.) on the internet, and these parasitic AI companies came in and built a fence around it and started charging admission while massively polluting our cities and towns and gobbling up resources at an incredible rate. Then, even if you pay them their admission fees, the information you get is at best 70% correct. Why would anyone celebrate that?
If a company unethically steals data and publishes a book with that data, we don’t say that there are legitimate concerns using books.
Off course we say its a legitimate concern if someone unethically steals data and publishes a book with that data… The point of scraping and stealing data without respecting its license, training their software with it and using it commercially (or even non commercially) without giving even credit and not respecting its original license?
the problem that people … have with AI isn’t with the technology itself nor its use, but the methods by which the companies “grow and harvest it”, so to speak.
No, its not just the methods. I am talking about deep problems with the technology, the methods how companies go with it, and then its usage from company and end user. Ai trained with those data and used makes it impossible to respect the original license, for the user. So in this case, the company using the Ai AND the user of the model violate licenses. Maybe. Maybe not, if the data allows it. But we don’t know and CANNOT check and know for sure, that’s the problem. Just because you don’t understand the issue or don’t care does not make it less of a problem.
As an example with software licenses like GPL requires that any derivative works using this source code (such as Linux or whatever my personal program is licensed under GPL) has to be GPL licenses (Open Source) too. That is not debatable. That is how this license work. Now the Ai can be trained on this data, using the source code, and then? See where this goes? That is only one generalized example.
The difference is, you know which book you have a concern with. Ai does not operate like that. With using Ai, you don’t know which book the data is from and it is not possible for the Ai to tell you that. There are conceptional problems with Ai. Your book example does not work here, because that works differently than Ai.
This is like saying that smoking weed is unethical because of the murderous actions of the cartels who grow and distribute it. You’re problem with AI isn’t with the technology itself nor its use.
I’m not sure about the weed argument, can’t say anything about it. But first, my personal problem with Ai doesn’t matter to this disussion, as I was talking about what people could have an issue with, not about my personal problems. Secondly farming data without respecting the original license, not even linking to the source where it got from is in fact a problem with “its use”. That is an unsolved problem lot of people just ignore, which does not make it less of a problem (it makes it a bigger problem). And that is just a few issues I listed, not even talking about the problems of generated content. There are legitimate concerns using Ai, no matter how you frame it, some just choose to ignore it.
Okay. Then the problem that people, not you particularly, have with AI isn’t with the technology itself nor its use, but the methods by which the companies “grow and harvest it”, so to speak.
Which I don’t even think is true. People have just lost their minds over this. If a company unethically steals data and publishes a book with that data, we don’t say that there are legitimate concerns using books.
We would be having a problem with it if every book was regurgitated stolen data. That’s kind of the core of how all of these generative AI models work to the point that their creators have argued in court that they wouldn’t exist at all if they weren’t allowed to steal all the training data.
It can be trained on the exact same free data that we can be trained on. Just like us, it doesn’t have to bypass paywalls. I think they’re arguing that if I can look at the Wikipedia and tell you that an avalanche in Pakistan killed ten climbers, then AI should be allowed to as well.
This is misleading and dishonest. Its not just looking up a single free information. Ai does more than that. Ai will not only scrape and copy the entire article, it will also scrape and copy the entirety of Wikipedia. There are licenses attacked to the Wikipedia article we have to follow to, if we scrape and use it for other purposes. It will also scrape the entire web. That is not the same as just looking up an Avalanche in Pakistan killed ten climbers.
This is weirdly accusatory. I’m not sitting here in a black cape and stovepipe top hat, twirling my mustache.
If this is the way we’re going to look at it then I guess we have to throw this technology out the window. It seems like a bad idea to do that though.
BTW so you agree with me then, because the only solution you see here to throw it away. Because what I described is not hallucinated by me, this is how its operate. It was to correct you for your wrong example with looking up a single information from Wikipedia.
Why? Besides that’s not possible in the near future, it could be made illegal and then suddenly the biggest companies (the driver behind it) would have to stop. This wouldn’t eliminate all Ai tools everywhere, so getting rid of it is probably not possible anymore.
But that wasn’t my suggestion either. And we don’t have to throw it away, we have to comply. It needs to change.
And this was my point from the very beginning. Using AI isn’t bad. The problems that people have with it has nothing to do with the technology itself.
I think you need to read up on the difference between “free as in beer” and “free” licenses. Everything on Wikipedia is copyrighted under a GNU license, which requires any use to give credit AND ALSO be licensed under the GPL. Which AI models do not (and cannot) do.
Also, just conceptually… We built a giant, free library (and museum and art gallery and etc.) on the internet, and these parasitic AI companies came in and built a fence around it and started charging admission while massively polluting our cities and towns and gobbling up resources at an incredible rate. Then, even if you pay them their admission fees, the information you get is at best 70% correct. Why would anyone celebrate that?
Off course we say its a legitimate concern if someone unethically steals data and publishes a book with that data… The point of scraping and stealing data without respecting its license, training their software with it and using it commercially (or even non commercially) without giving even credit and not respecting its original license?
No, its not just the methods. I am talking about deep problems with the technology, the methods how companies go with it, and then its usage from company and end user. Ai trained with those data and used makes it impossible to respect the original license, for the user. So in this case, the company using the Ai AND the user of the model violate licenses. Maybe. Maybe not, if the data allows it. But we don’t know and CANNOT check and know for sure, that’s the problem. Just because you don’t understand the issue or don’t care does not make it less of a problem.
As an example with software licenses like GPL requires that any derivative works using this source code (such as Linux or whatever my personal program is licensed under GPL) has to be GPL licenses (Open Source) too. That is not debatable. That is how this license work. Now the Ai can be trained on this data, using the source code, and then? See where this goes? That is only one generalized example.
We have a legitimate concern with that particular book, not all books.
The difference is, you know which book you have a concern with. Ai does not operate like that. With using Ai, you don’t know which book the data is from and it is not possible for the Ai to tell you that. There are conceptional problems with Ai. Your book example does not work here, because that works differently than Ai.