this post was submitted on 12 Oct 2024

222 points (95.5% liked)

Technology

59092 readers

6622 users here now

This is a most excellent place for technology news and articles.

Our Rules

Follow the lemmy.world rules.
Only tech related content.
Be excellent to each another!
Mod approved content bots can post up to 10 articles per day.
Threads asking for personal tech support may be deleted.
Politics threads may be removed.
No memes allowed as posts, OK to post as comments.
Only approved bots from the list below, to ask if your bot can be added please contact us.
Check for duplicates before posting, duplicates may be removed

Approved Bots

founded 1 year ago

MODERATORS

222

Reasoning failures highlighted by Apple research on LLMs (appleinsider.com)

submitted 3 weeks ago by Timely_Jellyfish_2077@programming.dev to c/technology@lemmy.world

59 comments fedilink hide all child comments

top 50 comments

sorted by: hot top controversial new old

[–] Technus@lemmy.zip 105 points 3 weeks ago (5 children)

These models are nothing more than glorified autocomplete algorithms parroting the responses to questions that already existed in their input.

They're completely incapable of critical thought or even basic reasoning. They only seem smart because people tend to ask the same stupid questions over and over.

If they receive an input that doesn't have a strong correlation to their training, they just output whatever bullshit comes close, whether it's true or not. Which makes them truly dangerous.

And I highly doubt that'll ever be fixed because the brainrotten corporate middle-manager types that insist on implementing this shit won't ever want their "state of the art AI chatbot" to answer a customer's question with "sorry, I don't know."

I can't wait for this stupid AI craze to eat its own tail.

[–] neshura@bookwormstory.social 28 points 3 weeks ago* (last edited 3 weeks ago) (5 children)

Last I checked (which was a while ago) "AI" still can't pass the most basic of tasks such as "show me a blank image"/"show me a pure white image". the LLM will output the most intense fever dream possible but never a simple rectangle filled with #fff coded pixels. I'm willing to debate the potentials of AI again once they manage to do that without those "benchmarks" getting special attention in the training data.

[–] GBU_28@lemm.ee 32 points 3 weeks ago* (last edited 3 weeks ago) (3 children)

Lol it do be that way

[–] GBU_28@lemm.ee 22 points 3 weeks ago

I will say the next attempt was interesting, but even less of a good try.

[–] Womble@lemmy.world 3 points 3 weeks ago* (last edited 3 weeks ago)

Thats actually quite interesting, you could make the argument that that is an image of "a pure white completely flat object with zero content", its just taken your description of what you want the image to be and given an image of an object that satisfies that.

[–] MP3Martin@programming.dev 2 points 1 week ago

Explanation here: https://youtu.be/NsM7nqvDNJI?t=13m45s

[–] Technus@lemmy.zip 19 points 3 weeks ago (1 children)

Problem is, AI companies think they could solve all the current problems with LLMs if they just had more data, so they buy or scrape it from everywhere they can.

That's why you hear every day about yet more and more social media companies penning deals with OpenAI. That, and greed, is why Reddit started charging out the ass for API access and killed off third-party apps, because those same APIs could also be used to easily scrape data for LLMs. Why give that data away for free when you can charge a premium for it? Forcing more users onto the official, ad-monetized apps was just a bonus.

[–] rottingleaf@lemmy.world 6 points 3 weeks ago* (last edited 3 weeks ago)

Yep. In cryptography there was a moment when cryptographers realized that the key must be secret, the message should be secret, but the rest of the system can not be secret. For the social purpose of refining said system. EDIT: And that these must be separate entities.

These guys basically use lots of data instead of algorithms. Like buying something with oil money instead of money made on construction.

I just want to see the moment when it all bursts. I'll be so gleeful. I'll go and buy an IPA and will laugh in every place in the Internet I'll see this discussed.

[–] gr3q@lemmy.ml 5 points 3 weeks ago* (last edited 3 weeks ago)

I tested chatgpt, it needed some nagging but it could do it. Needed the size, blank and white keywords.

Obviously a lot harder than it should be, but not impossible.

[–] rottingleaf@lemmy.world 3 points 3 weeks ago (1 children)

Because it's not AI, it's sophisticated pattern separation, recognition, lossy compression and extrapolation systems.

Artificial intelligence, like any intelligence, has goals and priorities. It has positive and negative reinforcements from real inputs.

Their AI will be possible when it'll be able to want something and decide something, with that moment based on entropy and not extrapolation.

[–] InternetPerson@lemmings.world 2 points 3 weeks ago (1 children)

Artificial intelligence, like any intelligence, has goals and priorities

No. Intelligence does not necessitate goals. You are able to understand math, letters, words, meaning of those without pursuing a specific goal.

Because it's not AI, it's sophisticated pattern separation, recognition, lossy compression and extrapolation systems.

And our brains work in a similar way.

load more comments (1 replies)

[–] theterrasque@infosec.pub 12 points 3 weeks ago (1 children)

I generally agree with your comment, but not on this part:

parroting the responses to questions that already existed in their input.

They're quite capable of following instructions over data where neither the instruction nor the data was anywhere in the training data.

They're completely incapable of critical thought or even basic reasoning.

Critical thought, generally no. Basic reasoning, that they're somewhat capable of. And chain of thought amplifies what little is there.

[–] AliasAKA@lemmy.world 2 points 3 weeks ago* (last edited 3 weeks ago)

I don’t believe this is quite right. They’re capable of following instructions that aren’t in their data but appear like things which were (that is, it can probabilistically interpolate between what it has seen in training and what you prompted it with — this is why prompting can be so important). Chain of thought is essentially automated prompt engineering; if it’s seen a similar process (eg from an online help forum or study materials) it can emulate that process with different keywords and phrases. The models themselves however are not able to perform a is to b therefore b is to a, arguably the cornerstone of symbolic reasoning. This is in part because it has no state model or true grounding, only probabilities you could observe a token given some context. So even with chain of thought, it is not reasoning, it’s just doing very fancy interpolation of the words and phrases used in the initial prompt to generate a prompt that is probably going to give a better answer, not because of reasoning, but because of a stochastic process.

[–] rottingleaf@lemmy.world 5 points 3 weeks ago

Synthesis versus generation. Yes.

And I highly doubt that’ll ever be fixed because the brainrotten corporate middle-manager types that insist on implementing this shit won’t ever want their “state of the art AI chatbot” to answer a customer’s question with “sorry, I don’t know.”

It's a tower of Babel IRL.

load more comments (2 replies)

[–] Lettuceeatlettuce@lemmy.ml 48 points 3 weeks ago (3 children)

Of course they don't, logical reasoning isn't just guessing a word or phrase that comes next.

As much as some of these tech bros want human thinking and creativity to be reducible to mere pattern recognition, it isn't, and it never will be.

But the corpos and Capitalists don't care, because their whole worldview is based in the idea that humans are only as valuable as the profitability they generate for a company.

They don't see any value in poetry, or philosophy, or literature, or historical analysis, or visual arts unless it can be patented, trademarked, copyrighted, and sold to consumers at a good markup.

As if the only difference between Van Goh's art and an LLM is the size of sample data and efficiency of an algorithm.

[–] leisesprecher@feddit.org 18 points 3 weeks ago (1 children)

You don't have to get all philosophical, since the value art is almost by definition debatable.

These models can't do basic logic. They already fail at this. And that's actually relevant to corpos if you can suddenly convince a chatbot to reduce your bill by 60% because bears don't eat mangos or some other nonsensical statement.

[–] Lettuceeatlettuce@lemmy.ml 7 points 3 weeks ago (4 children)

It's all connected, the reasons why it can't do basic logical reasoning are the same for why it can't replace human art.

It's because neither of those activities are mere pattern recognition and statistical inference, which is all LLMs will ever be.

load more comments (4 replies)

[–] rottingleaf@lemmy.world 2 points 3 weeks ago

I'm just thinking - 12 years ago there was a lot of talk of politicians and big corpo chiefs being replaceable with a shell script. As both a joke and an argument in favor of something requiring change.

One can say it was saying that these people are not needed - engineers can build their replacements.

In some sense AI is politicians and big bosses trying to build a replacement for engineers, using means available to these people.

Maybe they noticed, got pissed and are trying to enact revenge. Sort of a domain area war.

load more comments (1 replies)

[–] oakey66@lemmy.world 18 points 3 weeks ago (1 children)

I work for a consulting company and they're truly going off the deep end pushing consultants to sell this miracle solution. They are now doing weekly product demos and all of them are absolutely useless hype grifts. It's maddening.

[–] tempest@lemmy.ca 3 points 3 weeks ago (1 children)

So... Just another Tuesday for consulting then?

[–] oakey66@lemmy.world 2 points 3 weeks ago

No. In the non sales world, I've built some really cool solutions for clients.

[–] WalnutLum@lemmy.ml 17 points 3 weeks ago (1 children)

I still think it's better to refer to LLMs as "stochastic lexical indexes" than AI

[–] exocortex@discuss.tchncs.de 15 points 3 weeks ago (1 children)

AI in general is a shitty term. It's mostly PR. The Term "Intelligence" is very fuzzy and difficult to define - especially for people who are not in the field of machine learning.

[–] rottingleaf@lemmy.world 4 points 3 weeks ago (1 children)

So for those in ML it's easier?

[–] flying_sheep@lemmy.ml 1 points 3 weeks ago

No it's not, that's why some smart people are starring by defining a more interesting concept: educability.

[–] tal@lemmy.today 17 points 3 weeks ago* (last edited 3 weeks ago)

Apple's study proves that LLM-based AI models are flawed because they cannot reason

This really isn't a good title, I think. It was understood that LLM-based models don't reason, not on their own.

A better one would be that researchers at Apple proposed a metric that better accounts for reasoning capability, a better sort of "score" for an AI's capability.

[–] Timely_Jellyfish_2077@programming.dev 16 points 3 weeks ago

Research paper : https://arxiv.org/pdf/2410.05229

[–] MonkderVierte@lemmy.ml 16 points 3 weeks ago (1 children)

What, reasoning was an expected feature?

[–] Allonzee@lemmy.world 2 points 3 weeks ago (1 children)

https://www.cnet.com/tech/services-and-software/chatgpt-gets-new-o1-model-first-to-have-reasoning-for-hard-problems/

load more comments (1 replies)

[–] vonxylofon@lemmy.world 11 points 3 weeks ago (3 children)

I still fail to see how people expect LLMs to reason. It's like expecting a slice of pizza to reason. That's just not what it does.

Although Porsche managed to make a car with the engine in the most idiotic place win literally everything on Earth, so I guess I'm leaving a little possibility that the slice of pizza will outreason GPT 4.

[–] Michal@programming.dev 3 points 3 weeks ago

LLMs keep getting better at imitating humans thus for those who don't know how the technology works, it'll seem just like it thinks for itself.

load more comments (2 replies)

[–] rimu@piefed.social 8 points 3 weeks ago* (last edited 3 weeks ago) (2 children)

I tried it myself (changing the name and changing the values) but lost interest after 3 attempts and always getting the right answer:

https://chatgpt.com/share/670af65d-da08-800f-8ad4-c67782ee5477

https://chatgpt.com/share/670af672-45dc-800f-ac91-cc2811fa89c7

https://chatgpt.com/share/6709e80b-e5a8-800f-90d0-1af3418675ef

[–] A_A@lemmy.world 3 points 3 weeks ago (1 children)

Errors from your links like this :
Unable to load conversation 670a...6ed2c

[–] rimu@piefed.social 2 points 3 weeks ago (1 children)

Sorry! I've updated my links now.

[–] A_A@lemmy.world 3 points 3 weeks ago

"... So, Mary has 190 kiwifruit."
nice 😋🥝

[–] tinsuke@lemmy.world 3 points 3 weeks ago (1 children)

I wouldn't doubt that LLMs got some special input to deal with the specific examples of this paper, or similar enough.

load more comments (1 replies)

[–] Arn_Thor@feddit.uk 5 points 3 weeks ago (1 children)

Water is wet. More at 11

[–] Aatube@kbin.melroy.org 3 points 3 weeks ago (1 children)

Water isn’t wet, water wets things, and watered things are wet by the wet but the water ain’t wet as it simply causes wet and thus water isn’t truly wet as water is pure water and pure water isn’t wet and water is not wet and water isn’t wet it’s not wet it’s not wet it’s not dry it’s not wet and it’s not wet it is wet it’s wet and you can see it is wet but it doesn’t look like it it’s dry it’s just wet and it’s wet so I just need it and it’s wet it’s not like it’s dry it’s wet it’s wet so it’s not dry but it’s wet it’s not wet so it’s wet it’s not dry and it’s not dry it’s wet and I just want you know how it was just to be careful that I just don’t know what to say I don’t know what you can tell him I just don’t

[–] neshura@bookwormstory.social 6 points 3 weeks ago (1 children)

if water makes other things wet then most water is wet because it (usually) is surrounded by more water. qed

[–] LostXOR@fedia.io 2 points 3 weeks ago (2 children)

An alternative argument: Water generally makes things "wet" due to it forming hydrogen bonds with said things. Water also readily forms hydrogen bonds with itself. Therefore, water is wet.

[–] embed_me@programming.dev 5 points 3 weeks ago

AI could never

load more comments (1 replies)

[–] DarkCloud@lemmy.world 4 points 3 weeks ago (1 children)

Do we know how human brains reason? Not really... Do we have an abundance of long chains of reasoning we can use as training data?

...no.

So we don't have the training data to get language models to talk through their reasoning then, especially not in novel or personable ways.

But also - even if we did, that wouldn't produce 'thought' any more than a book about thought can produce thought.

Thinking is relational. It requires an internal self awareness. We can't discuss that in text so much that a book is suddenly conscious.

This is the idea that"Sentience can't come from semantics"... More is needed than that.

[–] A_A@lemmy.world 5 points 3 weeks ago

i like your comment here, just one reflection :

Thinking is relational, it requires an internal self awareness.

i think it's like the chicken and the egg : they both come together ... one could try to argue that self-awareness comes from thinking in the fashion of : "i think so i am"

[–] john117@mastodon.jmsquared.net 2 points 3 weeks ago

@Timely_Jellyfish_2077 interesting read, thanks for sharing

[–] lvxferre@mander.xyz 1 points 3 weeks ago* (last edited 3 weeks ago)

Here's a simple test showing lack of logic skills of LLM-based chatbots.

Pick some public figure (politician, celebrity, etc.), whose parents are known by name, but not themselves public figures.
Ask the bot of your choice "who is the [father|mother] of [public person]?", to check if the bot contains such piece of info.
If the bot contains such piece of info, start a new chat.
In the new chat, ask the opposite question - "who is the [son|daughter] of [parent mentioned in the previous answer]?". And watch the bot losing its shit.

I'll exemplify it with ChatGPT-4o (as provided by DDG) and Katy Perry (parents: Mary Christine and Maurice Hudson).

Note that step #3 is not optional. You must start a new chat; plenty bots are able to retrieve tokens from their previous output within the same chat, and that would stain the test.

Failure to consistently output correct information shows that those bots are unable to perform simple logic operations like "if A is the parent of B, then B is the child of A".

I'll also pre-emptively address some ad hoc idiocy that I've seen sealions lacking basic reading comprehension (i.e. the sort of people who claims that those systems are able to reason) using against this test:

"Ackshyually the bot is forgerring it and then reminring it. Just like hoominz" - cut off the crap.
"Ackshyually you wouldn't remember things from different conversations." - cut off the crap.
[Repeats the test while disingenuously = idiotically omitting step 3] - congrats for proving that there's a context window and nothing else, you muppet.
"You can't prove that it is not smart" - inversion of the burden of the proof. You can't prove that your mum didn't get syphilis by sharing a cactus-shaped dildo with Hitler.

load more comments