Final Unicourse'tan Çalış, Yüksek Notu Garantile!
Vizesine Unicourse'tan Çalış, Yüksek Notu Garantile!
AI fails badly in discussing books I am reading
-
I was recently reading Ben Aaronvitch’s ‘Rivers of London’ and I was struggling to follow the plot. So I engaged in discussions with both ChatGPT and Gemini and both really struggled to not make lots of mistakes. Both AIs would regularly give inaccurate info about major plot points, and when I questioned them about those issues, they would apologize profusely and then try to continue on. I basically had to coach them through to get them to remember important points of the plot! They also struggled to not reveal spoilers past a certain chapter. It was interesting to see the AIs repeatedly fail.
LLMs don't really know the subject matter they talk about in the same way a person does. Whatever you feed as input biases the responses towards and away from certain sequences of words. If you put in something about Star Wars, it'll weigh its output in favor of saying things like "Luke Skywalker" and "Darth Vader" and away from things like "The best way to use a crockpot is...". When you're trying to get it to recall facts about things that aren't extremely famous, it will jumble some parts or just make things up that sound plausible.
As an example, I liked to prompt them with "Tell me about 'A Valiant Effort (1989)'" as a test -- it doesn't actually exist, but they'd often tell me it was a painting or a movie or an obscure book with hallucinated details until very recently; newer models seem more likely to say "I don't know what that is" or similar, but that's not guaranteed.
To get better results, try feeding it a passage you'd like it to work on directly in the input instead of assuming it just knows the story already, and you'll get output weighted by the actual text (and whatever else you put in context).
-
I was recently reading Ben Aaronvitch’s ‘Rivers of London’ and I was struggling to follow the plot. So I engaged in discussions with both ChatGPT and Gemini and both really struggled to not make lots of mistakes. Both AIs would regularly give inaccurate info about major plot points, and when I questioned them about those issues, they would apologize profusely and then try to continue on. I basically had to coach them through to get them to remember important points of the plot! They also struggled to not reveal spoilers past a certain chapter. It was interesting to see the AIs repeatedly fail.
You run out of people to talk to or what?
-
You run out of people to talk to or what?
Ha ha. I'm experimenting with the technology.
-
LLMs don't really know the subject matter they talk about in the same way a person does. Whatever you feed as input biases the responses towards and away from certain sequences of words. If you put in something about Star Wars, it'll weigh its output in favor of saying things like "Luke Skywalker" and "Darth Vader" and away from things like "The best way to use a crockpot is...". When you're trying to get it to recall facts about things that aren't extremely famous, it will jumble some parts or just make things up that sound plausible.
As an example, I liked to prompt them with "Tell me about 'A Valiant Effort (1989)'" as a test -- it doesn't actually exist, but they'd often tell me it was a painting or a movie or an obscure book with hallucinated details until very recently; newer models seem more likely to say "I don't know what that is" or similar, but that's not guaranteed.
To get better results, try feeding it a passage you'd like it to work on directly in the input instead of assuming it just knows the story already, and you'll get output weighted by the actual text (and whatever else you put in context).
That's a funny way to test it with a fake movie!
-
Did the AI “fail,” or did you use AI in a suboptimal way?
If you had started the conversation by feeding in the actual text of the book, you would have very different results, because you would be grounding on context rather than asking the model to “figure it out somehow” based on training data.
Not sure what you were expecting, using AI is mostly about context management, and if you don’t set up the context correctly you’re asking for results similar to this.
It appeared it had access to the entire book - it linked me to a website that somehow had the text. I was able to coax better responses out of it when I corrected its mistakes about plot etc.
-
i meant book. i typed that on my phone without reading glasses
I didn't upload the book in this case. I could tell from the links that ChatGPT was using as references that it did have access to the full text through a website that had the text of the book, perhaps not legally.
-
I did have a much better discussion of 'Moby Dick' a while ago, but I imagine there is lots of info that AI has scraped from scholarly articles and the like about classics.
They don't do well with big sources like books unless they're discussed. If there's no online references or scholarly articles it won't know the contents of the book. Same with movies or music.
-
It appeared it had access to the entire book - it linked me to a website that somehow had the text. I was able to coax better responses out of it when I corrected its mistakes about plot etc.
It appeared it had access
Unless you provided it, the context window did not have the context. “Appeared to” is also not a valid way to evaluate technology.
-
Probability machine prone to hallucinations is how I call them.
hallucinations
No need to humanize them, they are Markov machines. That is all.
-
Ha ha. I'm experimenting with the technology.
Don't.
Next time just make a «[Discussion][Help] ’s Rivers of London — Ben Aaronvitch 9780575097568»🧵, and wait for everyone else interested in responding to do so.
Hello! It looks like you're interested in this conversation, but you don't have an account yet.
Getting fed up of having to scroll through the same posts each visit? When you register for an account, you'll always come back to exactly where you were before, and choose to be notified of new replies (either via email, or push notification). You'll also be able to save bookmarks and upvote posts to show your appreciation to other community members.
With your input, this post could be even better 💗
Kayıt Ol Giriş