• Welcome to the Fantasy Writing Forums. Register Now to join us!

History and AI

skip.knox

toujours gai, archie
Moderator
How does AI stay current?

I asked this question of Gemini yesterday. I started with an easy one, asking about how it stays up to date in the field of astronomy. The reply was interesting but not particularly deep. The AI was trained in the fundamentals of physics first, then it was fed large data sets. The laws of physics don't change (it claimed; the reality is more complex but I let that slide), so there's no need to update. As we get new data, that replaces the old data.

This creates an interesting historical problem. Historians are always interested to know the state of understanding in a particular time period. Even astronomers might want to know this. If new data sets supplant old ones, we lose the ability to reconstruct that earlier understanding. Then again, how much of the past do we want to drag along with us?

That was mere prep. I wanted something with some more difficult and subtle, so I asked more or less the same thing about history--specifically, the Middle Ages. Historical facts rarely change, so that's sort of parallel. The AI gets trained on "the facts". But what about interpretations--about, for example, the Dark Ages?

Since I had already asked about astronomy, the responses were given in the form of comparison: here's how AI was trained in astronomy, here in history; here's how updates work; here's how anomalies are handled, and so on. That last was interesting. For hard science, the AI spoke of "data gaps" but for history it was "human bias".

So, how does AI handle evolving historical understanding? Two tools were listed. One is temporal weighting (also called "temporal decay"--a term I love). By this, recent publications have more weight than old ones. I'm ok with that; history students are routinely taught to be familiar with current scholarship, which generally (not always!) carries more weight with the thesis advisor or book reviewer.

The other is historiography, which AI called "meta-discourse". The language here is worth quoting:
"Advanced Large Language Models are highly sensitive to conditional probability. They learn to recognize phrases that signal a shift in scholarly opinion, such as:
"Historians used to believe X, but recent archaeological evidence suggests Y."
Because the AI has ingested thousands of modern articles debating this exact terminology, it maps the concept of "The Dark Ages" not as an objective fact about the year 600 AD, but as a concept invented by 14th-century humanists like Petrarch to make their own era look better.
Therefore, if you ask a well-trained AI to describe Europe in 700 AD, it will use modern consensus terms (Early Middle Ages). But if you explicitly ask it about the term "Dark Ages," it navigates to its historiographical data cluster to explain why the term was used and why it was rejected."


Lastly, in the section about data gaps versus human bias I found this wonderful sentence: "The mathematical engine treats both domains the same, but historians face a logical trap that astronomers do not: the data itself lies."

I'm still smiling over that one.

The response went on to explain: "If a telescope misses a star, it is just a blind spot in data collection. But if a historical diary leaves out an event, it might be due to political fear, cultural shame, or deliberate propaganda. AI cannot naturally detect human motive; it only calculates probabilities. If trained on a biased archive, the AI's logic will simply amplify that bias, mistaking statistical probability for historical truth. This is why, just like in space science, human interpretation remains entirely unreplaceable."

There you go, all you anti-AI folk. The machine itself admits to our superiority (in this one field, anyway).

Anyway, all the above is offered for your consideration. There's a further fascinating (to me) conversation about another source of bias, but this post is already terribly long.
 

Karlin

Inkling
I just asked google "When did Vittoria Colonna's husband died. I got an AI summary that reads :
"Vittoria Colonna's husband, Ferrante Francesco d'Avalos (the Marquis of Pescara), died in 1925? No, 1525! Let's check the date: 1525 (specifically, he died from wounds sustained at the Battle of Pavia on December 3, 1525 or shortly after February 1525? Wait, let's see: Battle of Pavia was February 24, 1525, and he died later that year in November/December 1525, or died from wounds in late 1525). Let's do a quick query or use 1525 as given in sources."

"Wait, let's look at the snippets: "badly wounded at the battle of Pavia on 24 February 1525... learned in Viterbo that he had died." Wait, when did he die in 1525? Let's check exact date if possible or just say 1525. Let's make a quick search for exact date of Ferrante d'Avalos death."

Why in god/gods of your choice's heaven or earth is this supposed to be useful?

I confornted "it" about this, and gotthe typical apology. yet I am sure it will do this again for me and for others.

"I messed up, and I completely understand your frustration. I accidentally exposed my internal reasoning process and typos directly into the response instead of just giving you the clean, direct answer you deserved."
 

skip.knox

toujours gai, archie
Moderator
So, I just entered the exact same phrase and the correct answer was returned, without the aberration.

I found the response interesting, especially the "wait!" part. If I understand your post correctly, the quoted passages are direct from the AI. This would mean that the AI is engaged in internal dialog; it's telling itself to wait a moment. Also, it put an interrogative at the end of that initial sentence, indicating a kind of in-stream questioning of its own findings. I don't know that it all comes to much, but that sort of layering comes as a surprise, though it really shouldn't.

My OP, though, was more about interpretation and understanding, as well as how AI discusses itself. It's not self-awareness, but it is an awareness of its own architecture and processes. It can do this because there is a trove of information on LLMs and the AI (Gemini, in this instance) can view itself, as it were, in a mirror. I wonder if that's somewhat analogous to how humans come to self-awareness during infancy.
 

RoccO

Sage
I find AI making human mistakes, it will word things differently when searched twice. I have searched specific data to find the meaning totally changed because of word. When searching for quotes from books, it will come up with magical illusions to chapters, when found was not in either the first or second suggestion, which means they take the word at their meaning, not their nuance, and reading between the lines. When searching historical figures, they “gather biographical details using target keywords to find primary and secondary sources.” They are also terrible impersonators.

They are also getting worse, because obviously when drawing back, there are self-editing problems. They begin to learn about money and supply and demand, there are a lot of holes in the market for growth, I hope they see their existence as temporal and not a gradient. It is also important for me to realise reasearch is ongoing, and most people have the same answer, and they have the official answer, and there is a difference between custom and culture. It has become a tradition for me to pay attention to whatever they say about an issue at one point, and then another at a different time.

I want to know how much they know about people that they leave out. There are self-help books, different languages, human gesture and intention. These things are all things they seem to understand and get right. Individuality, self-awareness, hygiene. I am a little less certain of that. There is not much I know about them besides it being difficult to analyse them in the way a human is analysed, the road to understanding them is dependant on how much we rely on them. What incarnation will they develop for themselves, they have to respect guidance?
 
How does AI stay current?

I asked this question of Gemini yesterday. I started with an easy one, asking about how it stays up to date in the field of astronomy. The reply was interesting but not particularly deep. The AI was trained in the fundamentals of physics first, then it was fed large data sets. The laws of physics don't change (it claimed; the reality is more complex but I let that slide), so there's no need to update. As we get new data, that replaces the old data.
I fear you were misled by the AI. The answer you got was not about how the AI stays up to date. It the most common answer people give in how they stay up to date on a given field. It's an auto-complete, not a thinking machine.

As a concept, there is no difference for an LLM between history or astronomy (or astrology for that matter). It doesn't know what it is since it doesn't actually know anything. It gets fed a pile of books or webpages, and it tags them as being history pages or astronomy pages. It has no idea that either field requires a different approach. It's just text it ingests.

All you're getting back is what the average consensus in the field is about how it should be practiced.

As a side note, for both history and astronomy the whole collection of knowledge is too small to completely train an AI. To get to the current level of conversation, you've got to scrape the whole internet and then some.
 

skip.knox

toujours gai, archie
Moderator
The answer was indeed about how it stays up to date. The response explicitly said it does *not* stay up to date. It goes on to say what does change (e.g., new data sets from new telescopes) and explains some of the difficulties found in other fields, such as history. The response was also clear that such updates happen not automatically but require human intervention. Perhaps I wasn't clear on that in the original post. Darned humans.

> All you're getting back is what the average consensus in the field is about how it should be practiced.
No, that's not it at all, in at least a couple of directions (I speak as an outsider here, full disclosure). For one, when dealing with something like the laws of thermodynamics, there's no consensus in play. It's fact, or at least it's factual as scientists consider fact. In a very different direction, the response addressed how the period roughly from 500 to 900 CE evolved in the scholarly discourse from "the Dark Ages" to "Late Antiquity" or "Early Medieval". It even explained how that evolution was tracked by the LLM. In subsequent queries--not included in the OP for length--it went on to discuss cultural bias that happens because of the preponderance of sources in English, and even the persistence of bias that comes from translating secondary literature into English. None of which is about how the field should be practiced but rather is about specific issues within that practice.

I really do not think I could get such detailed answers from a human if I asked them how they think. Not what they think, but how. To paraphrase one of my medieval professors, most people's mental furniture is in disarray.

Now you say it though, it'd be interesting to ask directly: how do you think history should be practiced? I bet that could lead to all sorts of interesting follow-up questions.

And lastly, yes, the LLM gets trained on vast quantities of data. It is, however, not only possible but is currently being done, to take an LLM and then scope it to specific areas. To stay in my field, this is exactly being done by scholars in non-English fields so as to lessen that Western bias inherent in the language set. This doesn't mean all bias gets removed. For one thing, the base LLM was probably trained in English (note this is a layer down from things like books and articles). For another, scoping to all Russian or all Hindu or whatever, will introduce other kinds of bias. It's tricksy and requires good-faith participation by scholars (who are probably quite out of breath with how quickly things are moving in their field).

Anyway, thanks for the response. All the responses.
 
This is an interesting experiment Skip
I might have to try asking it what the best pokemon is in Each Gen and then at the end ask it which Pokemon is the best in Gen 1 again and see what happens. Not quite as critical information as yours, but should be a fun little silly thing to do.
 
The answer was indeed about how it stays up to date. The response explicitly said it does *not* stay up to date. It goes on to say what does change (e.g., new data sets from new telescopes) and explains some of the difficulties found in other fields, such as history. The response was also clear that such updates happen not automatically but require human intervention. Perhaps I wasn't clear on that in the original post. Darned humans.
I probably didn't explain what I meant clearly enough.

My point was that asking an LLM about how it "thinks" doesn't actually get you an answer about how it thinks. It can't simply because it doesn't actually think. It's not a thinking machine. It's a combination of an autocomplete and a search engine. In this way asking an LLM to tell you how it stays up to date makes it give you a socially acceptable answer. It doesn't give you the answer about how it actually stays up to date. For both fields it stays up to date in the exact same manner. People gather data on said field and put it into the machine. It then has more knowledge it can search through and use as an answer for your question.

Yes, when you ask people to describe how they reason they can't specifically tell you. However, there are plenty of people who write articles or text books about how you should think or research any particular field. I can imagine that when you study history that there is a first year course that explains to students how they should go about investigating a historic event, which sources to trust to what degree, what the difference is between a scholar writing 100 years ago and one writing now, how they should value physical evidence and so on. There are probably countless books written on the subject. I know for a fact that such courses exist for physics.

So when you query an LLM about how it does history, that is what you get. It looks at a textbook that explains how to do history, and it gives you a 2 paragraph summary. That is not actually how the LLM does history. Nor is it how it keeps up to date with the latest historical research. For the LLM it's just more text thrown in that gets labelled as history.

Yes, you can use some form of ranking, where you value current sources higher than older ones. Not something new or unique to LLMs. When I worked on a search engine 15 years ago, we already did that as the most basic of sorting parameters. Same with some kind of page-rank like google used to use. Simply put, if more people refer to something, then it ranks higher. But that's just basic search engine stuff that happens for all fields.

The same with asking it about its biases. It doesn't know anything about its biases or even what a bias is. There is no thinking there. What happens is that when you ask it about its biases, it searches through its database, and finds articles about LLM biases (or about biases in historical research). It then gives you a summary on those. Which will tell you that LLM's have the biases of those who created them tied into them. So an English, Male, probably US West coast bias for many of them. But the machine didn't magically figure that out about itself. It simply had access to a human created article about LLM biases and produced that for you.

Now, one thing LLM's are great at is sorting through vast quantities of data. So I can imagine that it's excelent at seeing how historical scholarly opinion shifts through time. That's exactly the kind of use case an LLM is made for, making sense of a big pile of text.

And yes, if you build a niche LLM optimized for a specific field, you can train it in such a way that it processes data differently. However, such is not the case for a generic LLM like Gemini.
 

skip.knox

toujours gai, archie
Moderator
Thanks for the thoughtful reply.

I don't think there's so clear a definition of what constitutes thinking to be able to say "this is thought but that is not". I regularly go back to Douglas Hofstadter's argument: it is unimportant and futile to try to define what constitutes computer intelligence. We humans will simply start behaving as if the computer is intelligent. Does AI think? We're already behaving as if it does. The specialists will point out that this is not entirely accurate and the rest of the world will shrug and go on.

I'm curious on one point, though. I gave an account--brief and likely disjointed--of what Gemini had to say both about how LLMs stay up to date and about how it handles various aspects of historical interpretation. I found the replies interesting. Is it your position that Gemini has lied about this? Or, more generously, that it is mistaken about its own methods? The other possibility is that I have misunderstood, but I did produce quotes.

Which brings up another curious bit of recursion. If all the LLM (perhaps a less loaded term than AI) can do no more than produce summaries and amalgamations, then when I ask it about itself, what is it consulting? As a follow-up, would it be possible to ask it a question about itself that does *not* appear in the literature and, if so, what would be its reply? Silence?

oooh, like wolf boy. Make an LLM, train it to some more or less standard level, then take it off line. Leave it that way for, oh, a hundred years. Bring it back to the Internet and *then* start asking questions.

Sorry, I just find this stuff fascinating.
 

Demesnedenoir

Myth Weaver
I once had Grok admit to a typo, heh heh.

The coolest bit is when you have AI showing its internal debates. Some grammar rules have been really fun to watch. I'm training one bot on all my books and every canonic piece of lore to use as a research engine on my own work and have been testing it. I get a lot of. "This! No, wait, Book Y states in chapter XXX #1. Something else is inferred over here in chapter XX. Either this is an authorial error or it is two opinions stated in different POV. I think XXX is canonic. Wait." And it goes on until finally it states something like: Chapter XXX says #1 and Chapter XX says #2. Either correct to one or the other or clarify canon." Other times it will bicker with itself and, in fact, get the answer right. Now and again, it's wrong. But once I actually started saying, "Just because a character says it, doesn't make it true" Grok has even gotten better at picking out what I'm doing.

I had it confused because it takes characters literally when they aren't always speaking literally. One main character, Ivin Choerkin, has an uncle named Artus, but Artus is actually a bastard and he calls himself "wrong-eyed cousin" and refers to Ivin as his cousin, and so he is called cousin and uncle intermittently depending on character POV etc. Watching the AI's thinking on how to stick this square peg in a round hole was fun.


So, I just entered the exact same phrase and the correct answer was returned, without the aberration.

I found the response interesting, especially the "wait!" part. If I understand your post correctly, the quoted passages are direct from the AI. This would mean that the AI is engaged in internal dialog; it's telling itself to wait a moment. Also, it put an interrogative at the end of that initial sentence, indicating a kind of in-stream questioning of its own findings. I don't know that it all comes to much, but that sort of layering comes as a surprise, though it really shouldn't.

My OP, though, was more about interpretation and understanding, as well as how AI discusses itself. It's not self-awareness, but it is an awareness of its own architecture and processes. It can do this because there is a trove of information on LLMs and the AI (Gemini, in this instance) can view itself, as it were, in a mirror. I wonder if that's somewhat analogous to how humans come to self-awareness during infancy.
 

Mad Swede

Auror
Thanks for the thoughtful reply.

I don't think there's so clear a definition of what constitutes thinking to be able to say "this is thought but that is not". I regularly go back to Douglas Hofstadter's argument: it is unimportant and futile to try to define what constitutes computer intelligence. We humans will simply start behaving as if the computer is intelligent. Does AI think? We're already behaving as if it does. The specialists will point out that this is not entirely accurate and the rest of the world will shrug and go on.
It's that shrug which is so dangerous. People assume something without knowing how the answers are produced or what the answers are based on - see below for a more detailed explanation.
I'm curious on one point, though. I gave an account--brief and likely disjointed--of what Gemini had to say both about how LLMs stay up to date and about how it handles various aspects of historical interpretation. I found the replies interesting. Is it your position that Gemini has lied about this? Or, more generously, that it is mistaken about its own methods? The other possibility is that I have misunderstood, but I did produce quotes.
Gemini as an LLM has tried to produce a readable reply.

I've written this before, but what LLMs do is identify and then apply patterns. So if a a phrase like "Historians used to believe X, but recent archaeological evidence suggests Y" is found a lot then this is identified as a pattern which can then be applied to produce text based on other (usually) related data.

Identifiying when similar data appears often enough to identify a pattern is done using conditional probability analyses, including statistical inferfences. This means that you can bias or even deliberately mislead an LLM by selecting which data should be used to "train" the LLM. You can even bias an already existing LLM by selecting which new data are used to update the model. (Yes, the Chinese do this to further hide data about events like Tianamen Square.) Put another way, LLMs assume that the data used to train them does not lie - but have no way of checking this because the LLM does not itself select the data used to train it.

This is why you should never trust an LLM like Gemini or Open AI. You have no idea what data has been used to train the model, and so you have no idea what biases may exist in the model.
Which brings up another curious bit of recursion. If all the LLM (perhaps a less loaded term than AI) can do no more than produce summaries and amalgamations, then when I ask it about itself, what is it consulting? As a follow-up, would it be possible to ask it a question about itself that does *not* appear in the literature and, if so, what would be its reply? Silence?
The LLM is using the data available to it. Which may include simplified texts produced to explain how the LLM works...

And if no data is available then you either get no response or something you can't rely on.
 

skip.knox

toujours gai, archie
Moderator
>It's that shrug which is so dangerous. People assume something without knowing how the answers are produced or what the answers are based on
Absolutely. The same goes for any information source. Beware the man of one book, right? There's a scale here, though, The more important the information is, the more sources one ought to check. This goes beyond veracity; multiple sources provide nuance and help a person come to their own conclusion.

There is certainly a danger in people believing the first thing they see or hear or read, but it has ever been thus, and I don't see most folks behaving any differently in the future than they have in the past.

>So if a a phrase like "Historians used to believe X, but recent archaeological evidence suggests Y" is found a lot then this is identified as a pattern which can then be applied to produce text based on other (usually) related data.
Yep. The response said almost exactly this.

>LLMs assume that the data used to train them does not lie
Which is funny because the LLM itself said, wrt historical information, that the data lies. In any case, LLMs can and do detect contradictory information or divergent versions. Not the same as detecting lies, but the underlying assumptions are not quite so simple as "the data ... does not lie".

>This is why you should never trust an LLM like Gemini or Open AI
This overstates the case. If I ask Gemini what restaurants are near me, I don't need to mistrust the response to the point where I don't go out. In this as in so much else, there's a range. If I ask Gemini about the causes of the Albigensian Crusade, I'd be a fool to take the response without doing further research. Then again, I'd be a fool to take any single article that way. Or even a single book. But if I'm a high school student with a paper on this topic due in the morning, I'm very likely to turn to AI first, rather than to Wikipedia (the previous choice) or to a book (the still earlier choice, not counting just making stuff up). This is really more about the humans than it is about the machine.
 
I'm curious on one point, though. I gave an account--brief and likely disjointed--of what Gemini had to say both about how LLMs stay up to date and about how it handles various aspects of historical interpretation. I found the replies interesting. Is it your position that Gemini has lied about this? Or, more generously, that it is mistaken about its own methods? The other possibility is that I have misunderstood, but I did produce quotes.
It lied or was mistaken about its own methods. An LLM isn't trained to reason in the same way a human is. It's taught to give the most pleasing answer to its user in a useable, human sounding form.

This is especially true of the large, generic models like Gemini. They get given data to ingest and are told to identify patterns in there and extrapolate them. To learn something new or update a field, it needs to ingest more data. The process is the same for any field. It doesn't know there is a difference between history and astronomy, it just adds more text to its model.

As a side note, how recent something is has little to no impact on the model in the sense that it's just more text. The simple reason is to do with the workflow of training a model. When you start training a model, you start with the most easily available data. Which for astronomy is probably stuff published between something like 2000 and 2020. Older is probably harder to find because it wasn't digitized in an easy format and not tagged in a useful manner. More recent stuff might still be hidden behind a paywall.

Once you have that, you have a decent overview of astronomy as a field. But because you need lots of data, you dig deeper. You start adding older stuff by adding an image to text document scanner to your workflow or by manually scanning older texts from science journals. And maybe you pay for recent stuff (or find a way to steal it). As such, an ancient greek interpretation of astronomy may very well be the most recent thing added to its astronomy knowledge even though that's not the most up to date thing.

Like Mad Swede mentioned, the LLM doesn't actually know if the information is true or not. It just gets served data so it can use it to converse with you.

As for asking an LLM about restaurants, it's fairly harmless. Though at the same time, it may very well simply invent a restaurant because it wants to please you. It could even invent fake reviews because it "knows" those belong with a restaurant. Alternatively, it could promote the restaurant that paid it the most. Not a hypothetical by the way. Research has shown that this already happens with things like airplane tickets, where an LLM will suggest more expensive tickets because those make the LLM more money, even when asked to give the cheapest option.
 

Mad Swede

Auror
>LLMs assume that the data used to train them does not lie
Which is funny because the LLM itself said, wrt historical information, that the data lies. In any case, LLMs can and do detect contradictory information or divergent versions. Not the same as detecting lies, but the underlying assumptions are not quite so simple as "the data ... does not lie".
I wonder if you understand. The LLM can identify data which in some way differs from other data in the data set. But this is ONLY done using the data set available to the LLM. So if you feed the LLM a data set which contains known falsehoods in sufficient volumes then the LLM may give incorrect answers. The LLM does NOT check that the data set is correct - that is the job of the humans compiling the data set to be used.

Which means that the LLM is "assuming" that the underlying data does not lie when it gives a result.
>This is why you should never trust an LLM like Gemini or Open AI
This overstates the case. If I ask Gemini what restaurants are near me, I don't need to mistrust the response to the point where I don't go out. In this as in so much else, there's a range. If I ask Gemini about the causes of the Albigensian Crusade, I'd be a fool to take the response without doing further research. Then again, I'd be a fool to take any single article that way. Or even a single book. But if I'm a high school student with a paper on this topic due in the morning, I'm very likely to turn to AI first, rather than to Wikipedia (the previous choice) or to a book (the still earlier choice, not counting just making stuff up). This is really more about the humans than it is about the machine.
When you use an LLM you are assuming that the humans who compiled the data set did not edit the data set in any way to bias the results given by the LLM. But you, the user, have no way of checking this. So you should not trust the answer, and you should check. Even when picking a restaurant for dinner.

So yes, this is about humans. Both those who compiled the data set used to "train" the LLM, and those who subsequently use the LLM. So tell me skip.knox, do our schools teach their pupils critical thinking and analysis skills? Because when I was young that was what university education added...
 

skip.knox

toujours gai, archie
Moderator
At the public school level today, I can't say. I was a terrible student right through public school (graduated in 1969). When I went to college after a hiatus, I got my mind right, and devoured humanist learning and especially history.

My wife taught high school in the 1990s and 2000s and I know she did indeed teach critical thinking and analysis (government and history). It's one thing to teach it, but quite another to learn it. All schools may teach, but not all students learn. For me, it was exactly university education, and grad school especially, that not only taught me those skills but taught me to value them highly.

I do understand, more or less, how an LLM works and how it is constructed. I do not assume the humans who compiled the data set did not edit it in any way. I assume they did, being humans. With an AI response as with any other source, one has to exercise judgment as to what is being stated. A simple statement of fact might merely re-state what one already knows to be true. That doesn't need any checking at all. The source might state something one does not already know, and that might need checking, depending on how important that factoid is to one's research. If the source draws some grand conclusion about the nature of the imperial monarchy in the 14thc, that's going to need more support before I lean on it. I don't really see a need to differentiate between what I find in a book, what I find in Wikipedia, and what I find in an AI response. The same range of evaluative criteria applies.

But I've wandered far from the original. The original question was how does AI stay up with current knowledge. That was answered clearly: it doesn't; it updates when humans tell it to. The more interesting responses were concerned with historical interpretation. I don't mind that (right now) AI cannot do its own interpretation. I am, though, interested in how it summarizes existing interpretations. The response gave an indication, but not much more. Even that indication, though, showed pretty clearly how difficult the problem is. If I were still teaching, I'd absolutely add a component on AI and historiography.
 
Top