A.I. implementation into Google is killing data accuracy for good

Understood. And I concur with much of what you have said. We agree that the placement of Gemini at the top of Google search returns is a poor concept. And we agree that AI assistants can be useful in specific environments. Where we may disagree is on the source of failures at the human/AI interface.

As you say, "the mechanic can choose to not use the assistant". But, in this instance, that assistant refuses to leave the yard. I do try to ignore Gemini's imperious emissions as I scroll past. But is simply ignoring something which, reputedly, requires 10 times the energy as a basic web search a responsible move?

When you said "the mechanic can choose ... to use the assistant by being more specific, I assume that you were still referring to Gemini. I see warnings about fact-checking AI chatbot results. The inference is that such systems should only be trusted to satisfy the most idle of curiosities. Is that really where we have collectively arrived at after four decades of the internet?

BTW: Just to show that this not all dump-on-Google, I just stumbled across a 'Grok Conversation' in Czech which references SPF. Granted, it can be difficult getting clear meanings out of multiple discussions. But, if a 15-year-old student made such sloppy assumptions in their essay, a high grade would be very unlikely.
 
When you said "the mechanic can choose ... to use the assistant by being more specific, I assume that you were still referring to Gemini. I see warnings about fact-checking AI chatbot results. The inference is that such systems should only be trusted to satisfy the most idle of curiosities.
That inference is not exactly what I was going for. Rather than the most idle of curiosities, it's instead the most mundane of tasks that you can throw to AI and expect it to do well in. You can also do a lot of much more complex tasks with sufficient set up.
Is that really where we have collectively arrived at after four decades of the internet?
No it's not and I definitely wouldn't draw the inference as that's not what I was trying to express. What I said was that there are different applications of AI useful in different situations. In a search bar, I 100% agree having a chatbot answer questions is shitty. As an assistant though, the same AI can be made to do mundane and time consuming tasks as well as complex ones that requires you to do sufficient set up.
BTW: Just to show that this not all dump-on-Google, I just stumbled across a 'Grok Conversation' in Czech which references SPF. Granted, it can be difficult getting clear meanings out of multiple discussions. But, if a 15-year-old student made such sloppy assumptions in their essay, a high grade would be very unlikely.
I think discourse ought to be dumping on google, or any brand of bad AI rather than dumping on AI as a whole.

It's unclear to me why seeing a grok convo referencing SPF is... bad without the original query but I'm assuming if your query was from authoritative stories and used to write an essay, then yes SPF would be a poor source of information.

This comparison is precisely the kind of misunderstanding that leads people to trash AI in ways that ... don't really make sense. AI was, and still is a tool built for great depth in narrow usecases or shallow depth in wide use cases and it's purpose isn't to be smarter than you. It's purpose was and still is the fast organization and retreival of information from massive corpuses. What it thinks it understands is, at its most fundamental level, a floating point measure of similarity between words, with algorithms built ontop of that fundamental fact to give it more (albiet still limited) hierarchical understanding. For most AI, even the "hierarchical techniques" here are heuristics programmed by the model's developer. Even in cases where AI is the one building these hierarchical connections and units of information organization, it's still limited both by the hardware and by the available algorithms. So with this understanding, comparing generic AI (and probably the most generalized model version at that) to a 15 year old child (or even a 5 year old child) is just not comparable. If you expect AI to have human like intelligence, ofcourse it'll fall short. It's not meant to be that nor can it be that even in the unforeseeable future.

However, with that understanding I highlighted, if you see it for what it is (a dynamically programmable tool) and leverage it in the correct ways and with the correct expectations, you'll be surprised at the kind of shit it can process and process better than humans can.

I've got two examples of a good usage of AI and a poor one.

I told GPT-5 to convert an F15E texture into the same color pallet as the B-21 ghost gray and to only process changes on the outer skin while leaving the rest of the internal components the same. This same process would have taken me days before and still generated a flat, de saturated livery, but instead, it achieved this without de saturating and losing the light variations across the textures. It's not obvious in the screenshot, but it's was a decent shortcut that shrunk my entire process from somewhere around a 2 - 3 days of tweaking down to less than 5 hours. I'm sure a more experienced modder than I could do the whole thing in less than 5 hours by hand, but at that point, they might even do it in less than 2 hours with some AI help.

And before anyone tells me I'm plagiarizing another artists work, I've never released this mod and I only mod for my own personal use.
1758072895682.png
This is a zoomed in view of one of the textures
1758073364243.png
I'm not great at photoshop, but I've always struggled with adjusting the underlying color without completely losing the color variation that makes a texture beautiful. With what I did before, the three different colors in the above screenshot would be three shades of the same color. I wouldn't expect AI to be able to make this without sufficient set up at all. But now that I set it up, I cranked out 7 of these textures in a matter of minutes and all that was left was to do some masking and remove / doctor up the places where AI missed or screwed up.

It'll never be something I can publish nor do I think it's as good as what the original artist produced, but it's been a huge boon to my personal enjoyment of video game mods.
---

My poor usage of AI example is that I've been playing around with a 3d model which I consider to be one of the most accurate of the F-22. I can't load the model into a 3d viewer because it's an .edm file, but what I can do is estimate coordinates, map them to a top down view and make the AI estimate things like the center of lift and mean aerodynamic chord. In order to conduct this mapping from a set of 3d coordinates to a top down view that has a different way of organizing axes, ChatGPT broke its damn head on trying to map these points onto the top down view. It took me literal days of trial and error to get it to finally work.

Other poor usages of AI was when I was referring to the J-36 6th gen fighter and chatGPT kept thinking I was talking about the J-35. These two planes are similar enough contextually that ChatGPT continually got it mixed up even after I corrected it.

I also tried to generate a design of what the F-47 might look like. I told it to draw me a tailess version and ChatGPT literally could not fathom what a tailless fighter looks like (it just kept giving me smaller and smaller horizontal tails, which began to look hilarious. Why? because by and large, taillessness in fighters is considered to be an extremely niche and small part of the corpus that ChatGPT likely was trained on and could see (and it doesn't help that there's any number of horrible speculation models as well as high quality ones that are 6th gens with tails). Had I built my own pipeline that specialized in aerospace engineering, I'd bet you that it would easily draw a fighter design that was tailless.

So yeah - AI and especially generic AI is narrowly good at a few things (and usually the more numeric your task is the better), mediocrely good at many easy things (like writing a script) and really bad at most things that are either obscure, features too much conflicting evidence, or highly correlated and while still being separate. Despite how much work that's been done since I left academia, these overall difficulties I've encountered with AI highly correspond to the issues that were being learned about and studied 7 years ago when I worked briefly in the area. Hopefully this illustrates what I mean when I say that AI is a dynamically programmable but specialized tool rather than some all knowing guru let a lone a replacement of a human assistant.
 
Last edited:
Yes, you're right about my aside ... I should have provided a link. It was:
-- https://x.com/i/grok/share/wgRUuQztcgIXGJCWVmir11Luj

The subject was the Austro-Hungarian designer, Eduard Zaparka. In the original SPF discussions, it is clear that Eduard Zaparka disappears from meagre available online Austrian records in the mid-1920s. By 1927, an Edward F. Zaparka (of later Zap flaps fame) turns up in the US. Grok Conversation connects the two names without question. And yet the lack of confirmation of a connection was a big part of those SPF discussions.

I mentioned that Grok Conversation because it illustrated the horses-for-courses argument. Whilst I was glad to see even vague references being provided by a chatbot, important nuances of the SPF discussions seemed to have been lost on Grok. That struck me as very odd. IIUC, Grok uses a Large Language Model in order to focus on conversational skills. That focus would explain its wordiness along with an apparent inability to research basic biographical details of its subject. And yet I could.

I'm not showboating here. I am a human with the same research limitations as a chatbot - ie: I have no access to books on the subject at hand; everything I glean about Zaparka came/comes from internet searches. Are Grok Conversation's seeming inability to parse SPF discussions an example of what was meant when you said: "For most AI, even the 'hierarchical techniques' here are heuristics programmed by the model's developer"?

Reworded for this specific case, was Grok Conversation's potential for true conversation limited by its developers? Or must we wait for future iterations capable of matching my hypothetical 15-year-old with a failing grade? Here, I realise that I may be demanding top performance from a toddler to justify its food bill. That would never occur with a human child but I don't believe that I am revealing an personal technophobic tendencies.

For me, that imaginary 15-year-old connects with your new example of amount of guidance required for ChatGPT to produce a reasonable illustration of a tailless fighter aircraft. Without assistance, both will have difficulty producing work with the smallest degree of originality (without the product being unintentionally hilarious). Personally, I faff about with a version of Photoshop from the last century. My software is old and my skills with it are limited ... but even I could produce a passable tailless fighter image (and with far less electricity consumed).

Going back to horses-for-courses, we have agreed on areas where AI can accomplish worthwhile tasks at speeds unimaginable with human specialists. And perhaps my sidebar has pounded flat the subject of 'general AI' and chatbots. If so, my apologies for perpetuating the very wordiness of which I had accused Grok Conversation.
 
Yes, you're right about my aside ... I should have provided a link. It was:
-- https://x.com/i/grok/share/wgRUuQztcgIXGJCWVmir11Luj

The subject was the Austro-Hungarian designer, Eduard Zaparka. In the original SPF discussions, it is clear that Eduard Zaparka disappears from meagre available online Austrian records in the mid-1920s. By 1927, an Edward F. Zaparka (of later Zap flaps fame) turns up in the US. Grok Conversation connects the two names without question. And yet the lack of confirmation of a connection was a big part of those SPF discussions.

I mentioned that Grok Conversation because it illustrated the horses-for-courses argument. Whilst I was glad to see even vague references being provided by a chatbot, important nuances of the SPF discussions seemed to have been lost on Grok. That struck me as very odd. IIUC, Grok uses a Large Language Model in order to focus on conversational skills. That focus would explain its wordiness along with an apparent inability to research basic biographical details of its subject.
I've never used grok nor do I plan to really, but conversational models tend to possess the least reasoning abilities while placing much higher focus, as you said, on conversation. Actual reasoning models before GPT 5 at least would just spit out factoids in bullet points and are much more to the point.

I typed in the following query to GPT-5 Thinking, which is not the conversational model of GPT. It is a model specifically made to conduct more reasoning in it's problem solving ( Do note - I had to translate your screenshot because I couldn't read the language. After I had ChatGPT translate it, I deleted the conversation and explicitly told it not use any other context besides what it can find on the internet). I'm not a paid advertiser for OpenAI. I just happen to use it a lot to study.

1758154734539.png

I didn't read through your translated post so I have no idea how well it did on the other facts, but at the very minimum this model acknowledged the two identities and that there are no definitive sources linking the two at all. The rest of the response also didn't pull anything from SPF. This isn't me advertising for OpenAI. I believe that higher reasoning models from any other company should be able to do the same.

In my experience though, even the thinking models do have problems summarizing and understanding correctly. We will get to why later.

Are Grok Conversation's seeming inability to parse SPF discussions an example of what was meant when you said: "For most AI, even the 'hierarchical techniques' here are heuristics programmed by the model's developer"?
In short, yes.

My research area when I was in school was on transformers - the very model family that gave birth to the GPT model.

In your vanilla neural network and even deeply layered neural networks, one of the most difficult problems was, in human terms, "forgetfulness". Without diving too deep into the technical details - all neural networks before transformers struggled with long term memory, because all models relies on the propagation of the calculated error back to dozens, hundreds, and even multiple thousands of layers, each of another hundred weights. If you think of a layer as a sieve or a filter, with every passing layer, the amount of material left to filter gets smaller. Your top level filters will catch all the flak while the bottom level filter's won't even touch what's being filtered. With every passing layer, the amount to update actually gets smaller and smaller until at a certain point, there's no reasonable updates made anymore - this means that beyond a certain layer of weights in the model, the rest of the model stops learning completely.

To combat this, we developed first the LSTM family of models. The researchers developing this model added gates, new connections between layers, and different representations of cell state that allowed us to learn what to forget, what to keep, and what to write to the cell state of the next time stamp. This wasn't enough, because even the best LSTMs could only understand something around 1000 word queries reliably. The downside of not being able to parallelize over multiple time steps meant that we still had immense difficulty learning word and sentence contexts. When Transformers came around, researchers introduced a mechanism into the neural network stack where evaluating a single words meaning would also account simultaneously all the other words or tokens in that sentences (attention mechanism). This further expanded the memory of a transformer to something like 4000 word queries, which is enough for most contexts, but remember that as you start adding more information frameworks on top, you have to merge those into your model as input as well.

Just in natural language models, every single major breakthrough was a learning and reasoning pathway that AI researchers borrowed from neurology, and manually implemented into a model. Just the implementation of a single technique (like attention for example) doesn't mean they implemented the intricacies of how the exact same technique might actually work in the brain (and we literally can't do that because we still don't understand our own brains well enough for that). So what you have then is essentially a biologically inspired analytical technique. This doesn't even cover other model types for different mediums (like CNNs for images).

There's a huge amount of uncertainty, imperfection, and flaw with how you integrate a new reasoning pathway into a neural network. You can very well implement what looks to be the correct pathway but if you chose to manipulate the wrong data with the path way or if you choose to manipulate the wrong quantity of data with the same pathway, it won't improve your model at all. This doesn't even go into superimposing huerisitics, and other analytical algorithms into a model yet. Indeed many specialized AIs are a mishmash monstrosity of different ML models, human developed heuristics, formal methods and entire reasoning frameworks trained and calibrated one ontop another in order to perform decently for a narrow set of tasks. Even transferring between different fields or different mediums require entirely different mismash of reasoning frameworks.

I hope you can see where this is going. On the surface - everything looks human-like and everything looks like it should be intuitive and easy to execute because these things come very naturally to ourselves. But under the surface? Far far far from that. This is why there are certain tasks that AI can do very well and do better than humans (think any task that is numeric in nature).
And yet I could.

I am a human with the same research limitations as a chatbot - ie: I have no access to books on the subject at hand; everything I glean about Zaparka came/comes from internet searches.

Reworded for this specific case, was Grok Conversation's potential for true conversation limited by its developers? Or must we wait for future iterations capable of matching my hypothetical 15-year-old with a failing grade? Here, I realise that I may be demanding top performance from a toddler to justify its food bill. That would never occur with a human child but I don't believe that I am revealing an personal technophobic tendencies.

For me, that imaginary 15-year-old connects with your new example of amount of guidance required for ChatGPT to produce a reasonable illustration of a tailless fighter aircraft. Without assistance, both will have difficulty producing work with the smallest degree of originality (without the product being unintentionally hilarious). Personally, I faff about with a version of Photoshop from the last century. My software is old and my skills with it are limited ... but even I could produce a passable tailless fighter image (and with far less electricity consumed).
I quoted these remarks wholesale because I think they highlight what I mean when I say that many people treat AI as they would a human and evaluate it as such. This is an erroneous evaluation that completely ignores the fundamental nature of AI. I'll try to use simplistic examples to highlight the difference. Pay major attention to 1 and 2. It's for those two reasons that trying to measure AI's reasoning ability to that of a human is not an apples to apples comparison. I would, instead, measure AI to other statistical methods in their ability to facilitate task. That's more fitting than comparing them to humans.
  1. For humans, we learn because our neurons automatically make connections with each other and those connections and their strengths activate our memories and knowhow. These connections are accumulated, automatically built, and automatically pruned as we study, practice and apply.
    • AI tries to mimic this behavior by introducing essentially information filters (gates) and randomized dropout of connections that let certain information through while not letting others through - and this is exactly why LSTMs were a huge leap forward. But notice - these algorithms in AI are not naturally developed by the itself. AI can't itself grow these techniques and apply them to itself. Even though there are indeed models that try to organically try to build these connections the following caveats still apply:
      • Remember what I said above about integrating new reasoning connections? Even if you've built them, if you use it on the wrong data or the wrong quantity of data, it would at best make no impact and at worse make your model drop in accuracy.
      • Also - remember what I said about memory and association? It's not about how much computer memory AI runs on. It's moreso about how far the gradient can travel that limits how much information AI can understand. Even in our current day and age, that gradient memory has a long long ways to go before it can even match a 15 year old.
  2. For humans, the vessels in our lives that capture meaning are words, with their inherent meanings given to us by years upon years of education, usage and reading. When I tell you "I feel ecstatic", you know what "ecstatic" refers to. You can feel what that word means.
    • For AI, their fundamental unit of meaning are numbers and numeric relationships. If I asked you to use a number to represent "ecstatic" you would look at me like I'm stupid. If I asked you to numerically quantify how similar the words "Zaparka" and "aircraft" are related, you'd look at me like I'm dumb. Yet - that's exactly how AI functions under the surface - even the natural language ones. The "closeness" and "associativity" of words are just statistical and numeric relationships.
    • For you the user, you are baffled by why AI seemingly can't do what 15 year old can, but under the surface, the way AI actually "thinks" about things are almost nothing like humans. There reasoning pathways (algorithms like attention, dropout, gated units etc) merely inspired by human thinking pathways.
    • Thankfully, it happens to be AI's proclivity for numeric relationships that make anything numeric a much much easier task for AI to be good at.
  3. For humans, when we think about a tailless aircraft, we might begin with the understanding that a fighter has engines, needs a good aspect ratio, or has a certain wing sweep. We set out to make a drawing of the fighter already with a logical framework precisely tailored to the topic at hand. So when we think about a fighter's tail, we know exactly what constitutes a tail. Not just that but we make connections to other large areas of context that is completely lost on AI - things like fuel capacity, weapon capacity, RCS reduction techniques etc. These things do not factor in gradually after hours upon hours iteration. They are instantaneous for us.
    • For AI, it understands images in terms of different (and limited) size and resolution of shapes and lines. In the same pipeline (in the case of a generic AI), you'll have a reasoning model that vaguely understands a fighter needs a cockpit, wings, tails, engines, and landing gear from all the training data it's ingested. However, I can gaurantee you that even the higher level reasoning models won't have developed the same hierarchical and categorical understanding that even someone like me with no aerospace background can have. It's concept of "tail" is a numerical association between different shapes, lines and maybe their relative positions taken from millions of pictures of training data. If your training data lacks pictures of tailess aircraft - and I'd argue that the proportion of pictures of tailed vs tailless aircraft is extremely lopsided - it's not going to have strong numerical associations of what constitutes a tailless aircraft.
I think AI devs, researchers already has a lot of trouble communicating what the words "reasoning" and "learning" mean in an AI context let alone the marketing team, who most definitely doesn't care. Of the points I named above, 1 and 2 (particularly 2) is precisely why expecting AI to be human is just the wrong approach to viewing AI. It's not human. It'll never be human. We, as users and as those seeking to harness AI, need to stop viewing AI in human lenses. The very things you expect to be easy for a human is not going to be easy for an AI. The reverse is also true.
 

Attachments

  • 1758160880171.png
    1758160880171.png
    100.9 KB · Views: 4
Last edited:
I'm not trying to be intentionally annoying but rather to demonstrate the difference in model usage as well as contextual usage. Fed exactly the same prompt GPT5-Thinking:

1758220652102.png

And here is Gemini (the actual webapp and also the simplest and fastest model) responding to the same prompt:
1758221331795.png

So ... I hope this illustrates what I mean when I say that certain implementations of AI are absolutely ass.

I regularly read that thread for news and it's extremely painful to read and pick out what's useful. I woulda thought that because this forum literally talks about every other aspect of military technology with a moderate to high level of technical engagement, people would be more open to investigating how the technical aspects of AI actually apply in military tech but ... that's either scattered, or drowned out by ridicule, snark and complaining with little other attempts made to actually understand the technical aspects AI.
 
Last edited:
You can say what you want but my (and indeed others) point was that people just taking the AI answer from a google search need to do so with caution as it is often wrong and sometimes grossly so. There is a growing laziness of people and they don't check sources or similar. This isn't just an AI issue. Look at Wikipedia for instance. I think it is getting worse with AI though and especially when it's outputs are put at the top of Google searches.
 
You can say what you want but my (and indeed others) point was that people just taking the AI answer from a google search need to do so with caution as it is often wrong and sometimes grossly so. There is a growing laziness of people and they don't check sources or similar. This isn't just an AI issue. Look at Wikipedia for instance. I think it is getting worse with AI though and especially when it's outputs are put at the top of Google searches.
That's fair and I don't disagree with that. IMO it's an irresponsible use of AI on the part of the developers. I don't even have any problems with the original topic here. I just have problems when it inevitably leads to people dismissing AI (like in the thread you linked).
 
That's fair and I don't disagree with that. IMO it's an irresponsible use of AI on the part of the developers. I don't even have any problems with the original topic here. I just have problems when it inevitably leads to people dismissing AI (like in the thread you linked).
I was just pointing to Reply #410.

I think AI has great potential if used appropriately. For instance, doing rapid analysis of data and pattern recognition - something that leads to all sorts of benefits including in medical arenas. I also think it has benefit in automating those tasks that are repetitive and not needing human involvement. Using it in other arenas though can sometimes be fraught with danger either through driving misleading/wrong understandings/conclusions etc (e.g. the cases in this thread with relying on Google AI search outcomes). It also will deprive people of opportunities to either learn themselves (e.g. students using Chat GPT or similar to do assignments), or to experience the joy/satisfaction of doing something for themselves (e.g. someone trying to use Chat GPT or similar to write a novel). Overall, it also risks people becoming mentally lazy which will have its own repercussions either in the dumbing down of society or worse, becoming targets to be fed misinformation. The latter is already a problem, but people risk making it even easier to dupe them. And those doing the 'duping' will not have good intentions...
 

Similar threads

Back
Top Bottom