“oversight” is a myth. I’m pretty sure you can’t oversight your way of plagiarism out of millions of training data sources that you don’t even have on your local disk for comparison and reference.
you are confusing two issues here: oversight for the output of an ANN can very much achieve good results. as i explained in my post.
oversight over the output won’t help you with copyright, water and energy consumption, slave labour and all the ethical issues further up the pipeline. but we don’t need ANNs for big corpos to do all these evil things. we need oversight and real consequences for those corpos - no matter what they produce.
ANNs as a technology are old and have not fundamentally changed since Alan Turing. Sure, we have iterated and improved. But the fundamentals are the same. LLMs just made that old tech quite popular recently and introduced a “line go up” race. i do not believe that we will gain significant improvements from simply feeding more stolen works to the machine. A fraction of the MNIST dataset is enough to train an ANN on a 20 years old laptop to recognise the digits 0-9 reliably. The technology is sound, limited in its usability and detached from the big corpos.
so why are you advocating for not avoiding all generative AI code then? you specifically cited Linus, who does as far as i know not ensure cleared up training data (what would that even be, CC0 only?) like you seem to be advocating for.
“oversight” is a myth. I’m pretty sure you can’t oversight your way of plagiarism out of millions of training data sources that you don’t even have on your local disk for comparison and reference.
(This isn’t legal advice. I’m not a lawyer.)
you are confusing two issues here: oversight for the output of an ANN can very much achieve good results. as i explained in my post.
oversight over the output won’t help you with copyright, water and energy consumption, slave labour and all the ethical issues further up the pipeline. but we don’t need ANNs for big corpos to do all these evil things. we need oversight and real consequences for those corpos - no matter what they produce.
ANNs as a technology are old and have not fundamentally changed since Alan Turing. Sure, we have iterated and improved. But the fundamentals are the same. LLMs just made that old tech quite popular recently and introduced a “line go up” race. i do not believe that we will gain significant improvements from simply feeding more stolen works to the machine. A fraction of the MNIST dataset is enough to train an ANN on a 20 years old laptop to recognise the digits 0-9 reliably. The technology is sound, limited in its usability and detached from the big corpos.
so why are you advocating for not avoiding all generative AI code then? you specifically cited Linus, who does as far as i know not ensure cleared up training data (what would that even be, CC0 only?) like you seem to be advocating for.