The (possible and degenerative) fate of AI models
June 16, 2024 · AI
The last year was a really big one for the AI. Nowadays, everyone is pouring money on these technologies, and soon data centers will be filled with power-hungry GPUs that will train and infer our models.
Sometimes, these models work like magic. But, what are the improvements we are going to see in the next few years.
Sad but true
Recently, my friend Diomidis Spinellis posted at X.

The paper he referenced is here.
So it seems that 52% percent of the ChatGPT answers around programming questions contained errors, and 39% was completelly overlooked by humans. So it seems that the current version of AI models, have a long way to go. They are impressive, but they have a long way to go. But this will improve over time right?
The feast of content
In order for these models to be trained, content is required, like source code, books, web pages, images, etc. And how is this content generated?
So far, by humans. But not for long, since content is monetized in the most aggressive ways, soon we will see (it is already happening) AI-generated content filling every website, like youtube, wikipedia, news sites, every single one of them. The question is, what is the error rate for this generated content?
I guess high. So, the AI models, will be trained in the future with generated content that will contain errors.
De-generative AI
These errors will create higher error rates in the answers, which will create higher error rates in the newer models and so on.
How can we exclude this content from the training? Somehow I guess, this will become more pressing concern in the future.
From a point on, nobody will know, what the correct answer really is, there will be no authoritative source for the stuff we are reading in the internet and in addition we will not exactly be able to easily identify erroneous answers.
The twilight zone
4 days later, Diomidis Spinellis wrote also in X.

I do not even know why they call it “hallucination”. :) It is a bug, reporting inaccurate content. Or maybe, we are using it, for the wrong purpose.
One problem at a time
It is true, that generative AI and LLMs in their current form solve one problem in a very efficient way. They can greatly understand human language (and many other form of input) and can generate if guided properly very human-like answers.
Maybe, we should stop treating them as databases and use them for that. I would be really happy, if in the near future, I could talk to my computer or phone, in my native tongue (Greek) and perform complex tasks, like in Star Trek.
But, as we said, one problem at the time. Till then, make sure that you really check your answers before publishing them to the internet, because soon or later they will come back to you.