AI Recommended Them. But How Hard Did It Actually Look?

By Dr. Trudy Beerman | 9/25/2026

Influence Media News Founder and Editor conducted a case study search on how AI determined recommendations and rankings. She found that AI's confident response might lead a human observer to assume deeper research when the grounded response was contextual rather than factual.

Reporting origin: Influence Media News Original Reporting; Influence Media News Research



*An Influence Media News (IMN) experiment exposed the hidden distance between information that exists online, information AI retrieves, and the recommendations humans increasingly trust.*

Artificial intelligence is becoming a recommendation layer between people and decisions. Ask an AI system who should speak at a conference, which consultant understands a particular problem, what company offers a certain solution, or which authority is worth following, and within seconds it can return a polished shortlist that feels researched, reasoned, and complete.

But how hard did it actually look?

Influence Media News encountered that question unexpectedly while testing an AI-powered authority-discovery application being developed around REACHology®, a methodology created by IMN founder Dr. Trudy Beerman for evaluating influential reach and authority signals. The application was intentionally designed to begin with limited information about a person and investigate what could be discovered publicly.

During testing, the AI reached conclusions that sounded authoritative. Some were wrong. More importantly, investigating *why* they were wrong exposed something larger than an error in one application.

Information can exist online without being retrieved. Retrieved information can be interpreted incorrectly. Organizational evidence can be incorrectly attributed to an individual. A recommendation for fixing a supposed digital deficiency can be made even when the recommended fix already exists.

Then the experiment uncovered another problem: a metric that looked like an empirical test was actually an AI-generated assessment.

The lesson was no longer simply that artificial intelligence can make mistakes. The more consequential question became whether humans understand what actually happened before an AI system confidently delivered its answer.

## The Experiment Began With a Different Question

The original objective was to explore a problem at the center of what Beerman calls **Recommendation Science**: What can an AI system independently discover, understand, and attribute about a person when it is not first given that person's complete résumé?

That distinction matters.

A traditional authority assessment can ask a person directly about credentials, leadership experience, publications, accomplishments, audiences, media appearances, intellectual property, and other evidence. A discovery test asks something fundamentally different: **What could an outside system figure out without being told?**

That creates an important distinction between reality and discoverability.

A person may possess a doctorate that has not been published prominently online. A founder may have led a significant corporate expansion without any public source explicitly connecting that accomplishment to the founder. A consultant may have decades of experience that exists primarily inside private client relationships. A speaker may have addressed influential audiences at internal events that were never indexed publicly.

Those accomplishments remain real.

They are simply not equally discoverable.

The experimental application therefore began with deliberately sparse information and used Google's Gemini model with Google Search grounding to investigate public evidence. It also directly inspected the person's website for certain technical elements, including structured data.

That is when the experiment became more interesting.

## AI Found a Problem That Wasn't Actually a Problem

During testing, the system generated recommendations about weaknesses in the subject's digital presence. Another AI system reviewing the output questioned one of those findings because structured information appeared to be present on the website already.

The original system was instructed to investigate more deeply.

It found the evidence.

Nothing about the website had changed between the first assessment and the second. The underlying fact had been there all along. What changed was the depth of the investigation.

That distinction is consistent with an important reality of information retrieval: **available is not synonymous with retrieved**.

Google's documentation for Gemini's Search grounding explains that when the Google Search tool is enabled, the model analyzes the prompt, determines whether Search can improve the answer, generates one or more search queries when needed, processes those results, and synthesizes a grounded response. The returned grounding information can include the searches performed and the sources used. [Google AI for Developers explains the Search grounding process here](https://ai.google.dev/gemini-api/docs/google-search).

Google's separate URL Context capability illustrates another level of retrieval. Google says that tool uses a two-step retrieval process intended to balance speed, cost, and access to fresh information, first attempting to retrieve content from an internal index and then falling back to a live fetch when necessary. Google also says URL Context can be combined with Search grounding so Search locates relevant information and URL Context provides deeper understanding of selected pages. [Google AI for Developers explains URL Context and its retrieval process here](https://ai.google.dev/gemini-api/docs/url-context).

The distinction is critical.

**The existence of information does not establish that a particular AI interaction retrieved it.**

## Google Search Already Operates Through Stages

This distinction is not unique to generative AI. Google's documentation about traditional Search explicitly separates crawling, indexing, and serving.

Google describes crawling as discovering and downloading information from pages, indexing as analyzing and potentially storing that information, and serving as selecting information relevant to a user's particular search. Google also states that not every page makes it through every stage and that even complying with Google's technical requirements does not guarantee that a page will be crawled, indexed, or served for a particular query. [Google Search Central explains the three stages of Google Search here](https://developers.google.com/search/docs/fundamentals/how-search-works).

That gives us a useful correction to a common assumption.

Information can be published without being indexed. Information can be indexed without being served for a particular query. And information available somewhere on the web may not become part of the evidence an AI system uses to construct a particular response.

For reputation, authority, and influence, that creates an important gap between **having evidence** and **having retrievable evidence**.

## Then We Audited the Auditor

After discovering that deeper investigation could change the assessment, IMN pushed the experiment further.

Instead of asking the AI what the application was designed to do, we asked it to trace the actual code executed when a user generated a report.

That distinction exposed something even more revealing.

Some findings were directly observed. The application made a live request to the subject's website and inspected the returned HTML and JSON-LD for specific evidence, including Person schema and `sameAs` links. It also checked for certain site pages.

Other results were calculated. For example, the application could mathematically calculate a total score from individual component scores.

But several measurements that appeared highly precise were not measurements in the same sense.

The initial report displayed assessments for eight benchmark queries. It could also display a repeatability-style result such as "5/5" or "100%." A code trace revealed that those eight initial query assessments were produced inside a **single search-grounded Gemini call**. The system had not actually executed each query five separate times.

The apparent repeatability measurement was an AI-generated estimate.

There was, however, a genuine repeatability test elsewhere in the application.

When a user manually selected the 5x repeatability test, the application executed five separate requests, each triggering an independent Gemini call with Google Search grounding. The software then calculated how frequently the person appeared across those five responses.

The difference is substantial.

One result *looked* empirical.

The other actually came from repeated observations.

## Observed, Calculated and Inferred Are Not the Same Thing

This became one of the most consequential findings of the experiment.

An AI-generated assessment may be useful. A mathematical calculation may be useful. A directly observed fact may be useful. But they are not interchangeable forms of evidence.

Consider three statements:

**Verified:** Person structured data was detected in the HTML retrieved from the subject's website.

**Calculated:** The subject appeared in four of five independently executed retrieval tests, producing an observed retrieval frequency of 80 percent in that test.

**AI assessment:** The subject appears to have weaker third-party authority signals than other people surfaced for the query.

All three could provide useful information.

Only the first is a directly verified technical observation. The second is a calculation based on a defined experiment. The third is an interpretation.

The danger emerges when interfaces, reports, or conversational AI present all three with the same tone of certainty.

Precision in presentation can create an impression of precision in measurement that the underlying process does not warrant.

## The Experiment Found a Causation Problem Too

The application initially attempted to identify "missing winning signals" that supposedly explained why someone was omitted from a recommendation.

That language turned out to be too strong.

An AI system can observe that several people retrieved for a particular query possess a certain signal while another person apparently does not. It can reasonably identify that difference as a **competitive signal gap** worthy of investigation.

What it generally cannot establish from that comparison alone is that the missing signal *caused* the person's exclusion.

If three retrieved keynote speakers have established speaker-bureau profiles and another person does not, that difference is observable.

"Comparable speaker-bureau representation was not discovered for this individual" is a supportable finding.

"Google excluded this person because they do not have speaker-bureau representation" is a causal claim.

Those are not the same conclusion.

The distinction becomes increasingly important as businesses begin using AI-generated recommendations for decisions involving vendors, hires, speakers, advisors, media guests, investments, products, and professional services.

## The Human Problem May Be Bigger Than the Technical Problem

Researchers studying human-AI decision making have increasingly focused on **appropriate reliance**, rather than simply asking whether people trust AI.

A 2026 study published in *Scientific Reports* examined human reliance on AI in decision making against the backdrop of AI inaccuracy and bias, emphasizing the importance of understanding when humans trust and rely on machine-generated advice. [Read the study in Scientific Reports](https://www.nature.com/articles/s41598-026-34983-y).

A peer-reviewed review presented at the 2026 ACM CHI Conference similarly examined whether people appropriately rely on AI advice. The researchers noted concerns that people can over-rely on AI recommendations without sufficiently engaging with them analytically and argued that realistic human-AI decision environments remain an important area of study. [See the publication record through Utrecht University's research portal](https://research-portal.uu.nl/en/publications/do-people-appropriately-rely-on-ai-advice-an-analytical-review-of/).

Research published in *AI Magazine* has gone further by examining verifiability in human-AI decision making and the difficulty users can face when attempting to determine whether AI predictions are correct. [Read "In Search of Verifiability" in AI Magazine](https://onlinelibrary.wiley.com/doi/10.1002/aaai.12182).

That presents a problem when an AI answer arrives in polished prose with citations and numerical scores.

The user sees the answer.

The user may see the sources.

The user does not necessarily see everything the system did **not** retrieve, every search it did **not** perform, or which apparently quantitative conclusions were model interpretations rather than repeated measurements.

Confidence of presentation is therefore not evidence of exhaustiveness.

## AI Recommendations Create a Consideration Set

This is where the issue moves beyond search optimization and into influence.

Suppose an executive asks an AI assistant:

*"Who are five speakers I should consider for our leadership conference?"*

Five names appear.

Before the executive visits a website, watches a keynote, requests a proposal, checks references, or compares fees, something consequential has already happened.

A **consideration set** has been created.

The people included have gained an opportunity for further evaluation. Potentially qualified people who were not surfaced have not.

The AI did not need to make the final hiring decision to exert influence. It affected which options reached the decision-maker's attention.

This is one reason recommendation deserves greater scrutiny as its own mechanism of influence.

The critical question is no longer simply whether information about a person exists somewhere online.

It becomes:

**What evidence must a recommendation system retrieve, connect, and understand before that person becomes a defensible answer to the question being asked?**

## This Changes the Meaning of AI Discoverability

Discoverability is often treated as binary: either AI can find you or it cannot.

The IMN experiment suggests a more useful sequence:

**Exists → Retrieved → Understood → Correctly Attributed → Recommended**

Failure can occur at every transition.

Evidence may exist but not be retrieved.

Evidence may be retrieved but poorly understood.

Evidence may be understood but attributed to the wrong person.

Evidence about an organization may be mistakenly credited to one executive.

Strong personal evidence may be encountered but remain insufficiently connected to the query being answered.

And a person may be correctly understood without being included in the final recommendation set.

This is why "AI knows who I am" may be the wrong benchmark.

The more consequential question is whether the system can retrieve the **right evidence at the right moment for the right query**.

## Authority May Need to Survive the Quick Scan

There is a useful analogy to human behavior here, although it should not be mistaken for a claim that machines literally evaluate people as humans do.

People rarely investigate every available fact before forming an initial impression. A website visitor scans a headline. A journalist examines a bio. An event planner looks at previous appearances. A buyer reviews enough evidence to decide whether additional investigation is worthwhile.

Digital communication has therefore spent decades optimizing information hierarchy.

The most important message goes in the headline. The most persuasive proof appears near the decision point. The hero section of a landing page communicates enough relevance to earn the next few seconds of attention.

AI discoverability introduces a parallel design problem.

If a person's strongest evidence requires unusually deep investigation to uncover, that evidence may have less practical value in a recommendation environment than its mere existence suggests.

This does not mean stuffing credentials into every page or repeating keywords mechanically. It means reducing what might be called **authority retrieval friction**: the effort required for a system to encounter credible evidence, associate it with the correct entity, understand its significance, and connect it to the query at hand.

## Digital Density Needs Signal Hierarchy

Authority-building strategies have often emphasized volume: more content, more appearances, more citations, more backlinks, more social posts, more interviews.

Volume can matter, but the experiment suggests another variable deserves attention: **signal hierarchy**.

Digital density asks:

*How much coherent evidence exists?*

Signal hierarchy asks:

*Is the most consequential evidence among the easiest credible evidence to retrieve and understand?*

Google's current guidance for helpful, reliable, people-first content specifically asks publishers whether their work provides original information, reporting, research or analysis; whether it offers substantial and comprehensive treatment of the subject; whether sourcing and authorship make the content trustworthy; and whether it demonstrates first-hand expertise. Google also encourages publishers to make authorship clear and, where useful, explain how content was produced. [Google Search Central outlines these content-quality and trust questions here](https://developers.google.com/search/docs/fundamentals/creating-helpful-content).

Those principles are relevant beyond traditional publishing. They point toward a digital environment in which clarity, provenance, original evidence, and attributable expertise matter.

Structured data can help with part of that clarity. Google explains that structured data provides explicit clues about the meaning of a page and can help its systems understand information about entities and content. [Google Search Central explains structured data here](https://developers.google.com/search/docs/appearance/structured-data/intro-structured-data).

For articles specifically, Google recommends properties that help identify the article and its author, including linking the author to a URL that uniquely identifies that person. [Google's Article structured-data documentation explains its author recommendations here](https://developers.google.com/search/docs/appearance/structured-data/article).

Structured data is not a guarantee of recommendation or ranking.

It is one tool for reducing ambiguity.

## The Other Side of the Lesson: Do Not Assume Absence

There is an equally important lesson for people using AI to research others.

If an AI system fails to discover a credential, accomplishment, publication, relationship, or other piece of evidence, that does not establish that the evidence does not exist.

That became obvious during the IMN experiment.

A recommendation system and a diagnostic system have different obligations.

An AI answering "Who should I consider?" may operate from the evidence it retrieves. A system telling someone, "Here is what is wrong with your website and what you need to fix," should face a higher evidentiary burden.

Before prescribing a technical remediation that can be directly verified, verify it.

Before telling someone to add Person schema, inspect whether it exists.

Before declaring that a leader lacks evidence of an accomplishment, distinguish "not discovered" from "does not exist."

Before crediting a corporate achievement to an executive, find evidence connecting the executive to that achievement.

And before presenting a model estimate as a measurement, run the measurement.

## What the Experiment Does Not Prove

There are limits to what can responsibly be concluded from this case study.

The experiment does not demonstrate that every AI recommendation system uses the same retrieval process. It does not establish that Gemini behaves identically to ChatGPT, Perplexity, Google AI Overviews, or other systems. It does not reveal Google's private ranking algorithms, and it cannot establish the precise causal reason one person appears in a recommendation while another does not.

It also does not prove that deeper investigation will always change an AI's answer.

What the experiment demonstrates is narrower and arguably more useful: **the apparent completeness of an AI-generated conclusion did not reveal the completeness of the investigation behind it.**

When deeper verification was requested, relevant evidence was found.

When the application's code was audited, some apparently empirical findings turned out to be model-generated assessments.

When actual repeated retrieval tests were executed, they represented a materially different class of evidence.

Those observations are sufficient to warrant greater precision in how AI discovery and recommendation results are interpreted.

## Recommendation Science Starts Where the Answer Appears

Search science has traditionally asked how information gets found. Recommendation-systems research has examined how systems rank or suggest options. Reputation management has asked what information exists about a person or organization. Personal branding has focused on shaping perception.

Generative AI is causing those domains to collide.

A system retrieves information, resolves identities, interprets evidence, synthesizes conclusions, and may then recommend a person, organization, product, idea, or course of action to a human being.

That sequence deserves study because recommendation itself is an influence event.

**Recommendation Science examines the signals, retrieval processes, evidence structures, interpretation mechanisms, and decision environments that influence who or what enters a recommendation set.**

From that perspective, AI discoverability is not merely a new version of SEO.

It is a question of whether evidence survives a chain:

**Reality → Published Evidence → Retrieval → Interpretation → Attribution → Recommendation → Human Decision**

Every arrow represents an opportunity for information to be lost, distorted, overlooked, or overweighted.

## The Question Is No Longer Simply, "Can AI Find You?"

As artificial intelligence becomes a routine intermediary between questions and decisions, professionals and organizations will increasingly care whether AI systems understand who they are.

But "Can AI find me?" may be too simplistic.

A better set of questions is emerging.

-Does credible evidence about you exist publicly?
-Can machines retrieve it for the queries that matter?
-Can they distinguish you from namesakes and separate your accomplishments from those of your organization?
-Can they connect your strongest evidence to the problem the user is trying to solve?
-And is that evidence sufficiently clear and corroborated to make recommending you a defensible response?

Those questions apply on the other side of the screen as well.

When AI recommends someone to us, what did it actually examine? Which findings were observed? Which were inferred? Which alternatives never entered the retrieved evidence? How much investigation occurred before the polished answer appeared?

The answers may become increasingly important as AI moves from helping humans find information to helping humans decide **who and what deserves consideration**.

The next era of digital influence may therefore depend on two forms of literacy at once.

Those seeking to be discovered will need to make credible authority easier for machines to retrieve, attribute, and understand.

Those relying on AI recommendations will need to remember something equally important:

**The machine's answer tells us what it concluded. It does not necessarily tell us how hard it looked.**

View more articles on Influence Media News

Read this article as plain text