NetContentSEO AI Labs: Testing What AI Systems Actually Understand
There is a lot of discussion about AI visibility, GEO and how brands will appear inside answers generated by ChatGPT, Gemini, Perplexity and other AI systems. The problem is that much of the conversation is still based on assumptions.
That is why NetContentSEO has started developing its AI Labs as a place for practical experiments rather than another collection of predictions about where search might be heading.
The basic idea is deliberately simple: create a reproducible test, give the same input to different AI systems and compare what actually happens.
The experiments include major commercial models such as ChatGPT, Gemini, Grok and Perplexity, but also smaller local models running without the same retrieval capabilities. The differences between them can sometimes be more interesting than the expected result.
One recent experiment, for example, tested whether an LLM could detect false information hidden inside an otherwise accurate technical article. Three fabricated or technically incorrect claims were inserted into a text about retrieval, RAG and embeddings, without telling the models how many errors they should expect to find.
Several systems identified all three underlying traps. A small local model produced a much stranger result: when it encountered a completely invented AI industry protocol, it attempted to correct the claim by inventing a different history for the same nonexistent protocol.
The experiment therefore became less about whether AI could spot a lie and more about what happens when a model encounters information that sounds plausible but cannot be reliably reconstructed.
The complete test, including the original prompt, test article and individual model responses, is available in the NetContentSEO AI Labs:
Can an LLM Detect a Lie Hidden Inside an Otherwise Accurate Article?
https://netcontentseo.net/article/can-an-llm-detect-a-lie-hidden-inside-an-otherwise-accurate-article-20-405
Other experiments have looked at different problems, including how AI systems debug the same broken PHP application and how different models reconstruct information about an author, a brand or an idea from the information available to them.
These are small experiments, not scientific benchmarks, and NetContentSEO deliberately treats them that way. Five or seven model responses cannot establish a universal rule about how LLMs behave. They can, however, expose behaviours worth investigating further.
That distinction matters.
AI visibility is often reduced to a new version of rank tracking: ask hundreds of questions, count brand mentions and produce a visibility score. Those measurements can be useful, but they only describe part of what is happening.
Being mentioned does not necessarily mean being understood correctly.
A system might retrieve the right source but associate an idea with the wrong entity. It might recognize a brand but describe products that do not exist. It might find accurate information and then reconstruct the relationship between those facts incorrectly.
This is the territory AI Labs is intended to explore.
The goal isn't to prove that one AI is better than another. It is to observe how retrieval, reconstruction, attribution and hallucination behave under controlled tests and document the results publicly.
As AI systems increasingly sit between publishers and their audiences, understanding those behaviours may become just as important as understanding how a traditional search engine ranks a webpage.
More experiments are available at:
NetContentSEO AI Labs
https://netcontentseo.net/category/ai-labs