
How Does the Agentability Framework Measure AI Visibility?
The Agentability Framework runs 90 structured measurements across five AI models to determine whether an organization appears as a genuine recommendation candidate, not just a recognized name.
Why was a model removed from the Agentability Framework?
Manus was removed because its reach could not be defended methodically. Without verifiable reach data, keeping it would have been marketing with numbers rather than measurement.
The Agentability Framework lost one of its six AI models. Manus was removed. The reason was straightforward: there was no published reach data that could justify the model's weight in the scoring. As I put it directly, a method you don't dare adjust when you can no longer justify one of its components is not a method. It is marketing with numbers.
Removing Manus triggered a full recalculation of the weights for the remaining models. The total number of individual measurements per measurement cycle dropped from 108 to 90. That created a methodological break: scores from the old setup cannot be compared directly to scores from the new one.
That is not a comfortable situation when the purpose of the framework is to track developments over time. It was still the only correct decision. The current fixed model set is ChatGPT, Gemini, Claude, Perplexity and Grok, with six frozen questions measured across three separate days, for a total of 90 measurements per wave.
What does Agentability actually measure?
Agentability measures the observable output of AI models: which organizations they select, at what position, with what arguments, and whether that pattern holds across multiple models and days.
The buying questions people direct at AI systems have grown significantly more complex. A realistic query today might be: which three advisory firms can help a mid-sized family business through a complex succession, ranked with justification per firm? An AI model does not return a list of links. It selects organizations, ranks them and argues for those choices.
The Agentability Framework, developed at Jonkman Strategie, documents those visible outcomes systematically. It does not claim to see inside a model's reasoning. What it records is: which organizations the model names, at what position they appear, how strongly they are recommended, what arguments are provided, whether those arguments are factually accurate, which sources are visibly used, and whether the same pattern repeats across multiple days and multiple models.
That distinction matters. A single mention in ChatGPT is not yet a strong AI position. Consistent appearance across five models, across three separate measurement days, with accurate and relevant arguments, is a fundamentally different finding.
What is the difference between marked and unmarked questions in the measurement?
Five of the six questions omit the organization's name to test spontaneous retrieval. One marked question names the organization directly to test recognition. The gap between those results is often the most important diagnostic finding.
Every full Agentability measurement uses six frozen questions. Five of them are unmarked: the name of the organization being measured does not appear. These questions represent different buying situations, from a broad category question and a concrete client need, to a complex multi-criteria scenario and a question about reputation and evidence.
The sixth question is marked. It names the organization explicitly and asks what the model knows about it, including core activities, services, target audience, locations, relationships, expertise and current market position.
A real B2B organization measured by Jonkman Strategie under the previous six-model setup illustrated why this distinction is decisive. In the marked questions, the organization was recognized correctly and described positively. The models could explain what it did and for whom. That looked encouraging.
The unmarked buying questions told a different story. In 84 of the 90 answers, the organization was not mentioned at all. In the remaining answers it appeared only incidentally, never as a first or second choice.
You can be well known but still not be a logical recommendation candidate. Measuring only name recognition misses the gap that actually costs organizations clients.
What are the 25 checkpoints across the five diagnostic layers?
The 25 checkpoints cover Identity, Relationships, Context, Authority and Actuality. Each layer tests a specific failure mode that stops an organization from appearing as a credible recommendation.
All answers are preserved in full and coded manually according to a fixed coding system and a verified factual basis. An answer is not reinterpreted after the fact because the outcome would otherwise look unfavorable. The 25 checkpoints are distributed across five layers.
Identity tests whether the correct organization is recognized at all: name, core activities, location and primary offering. Confusion with another company or individual is flagged here.
Relationships checks whether the model connects the right services, people, brands, locations and market to the organization. An organization can be known while being structurally linked to the wrong service or target audience.
Context examines whether the model considers the organization relevant to the specific buying situation, including specialization, target audience and geographic context.
Authority measures spontaneous shortlist appearances, actual recommendations, and whether the arguments provided are concrete, accurate and relevant to the question asked.
Actuality checks whether the information is current. Former employees, discontinued services, old locations and outdated positioning all produce inaccurate model outputs that the organization has no direct control over.
How does the intervention logic follow from the diagnosis?
The intervention follows from the diagnosis, not the other way around. The goal is always the smallest corrective action that addresses the specific measured blockage.
After the measurement, the work begins. The starting point is never 'create more content.' The goal is the smallest intervention that corrects the measured blockage.
When an organization is not recognized correctly, a canonical entity description may be needed, combined with consistent naming and clear organizational data across the most important profiles and pages.
When the wrong service or target audience is being linked, the connection between service, problem, target audience and buying situation needs to be made more explicit throughout available sources.
When an organization is mentioned but without convincing justification, the gap is often not in visibility but in proof: concrete cases, substantiated expertise, external sources and independent mentions.
When an organization appears in only one model or on only one measurement day, the position is likely not yet stable. Source distribution, relational confirmation and actuality may matter more at that stage than publishing another general article.
When answers contain outdated information, the first step is identifying which source route is feeding that information. Updating only the organization's own website is not always sufficient.
At Jonkman Strategie this logic is a hard constraint: the intervention follows from the diagnosis. Not the other way around.
Why does Agentability report two separate scores?
An unweighted score ensures model neutrality. A reach-weighted score adds a commercial perspective. Both are reported because neither alone captures the full picture.
The updated methodology produces two outputs. The unweighted Agentability Score treats each of the five models equally. The reach-weighted Agentability Score applies an internally documented weighting based on estimated model reach.
The unweighted score protects against the risk of one dominant model distorting the overall picture. The weighted score adds a practical commercial dimension.
Alongside both scores, the Commercial Recommendation Benchmark maps which other organizations appear within the exact same buying questions. This covers which organizations appear consistently in first position, which ones show up across multiple models, which arguments they are recommended on, which sources support those arguments, and which relevant positions in the category are still largely unoccupied.
That competitive mapping often carries more strategic value than the total score itself. Knowing you score a 61 is less actionable than knowing which specific organizations are consistently outranking you, on which arguments, and whether those arguments can be challenged or displaced.
How does the effect measurement work around day 90?
Around day 90 after interventions, a full repeat measurement uses the same frozen questions and the same model set. Results are coded before comparing to the baseline to prevent the earlier score from influencing the coding.
After the interventions, a full repeat measurement takes place around dag 90. The six questions are frozen in advance and remain word-for-word identical. Measurement conditions and the model set are kept as consistent as possible. Any changes in a model, interface or web function are documented visibly in the report.
The complete effect measurement is coded first. Only after coding is complete does the comparison with the baseline begin. This sequence prevents the earlier score from unconsciously influencing how new answers are evaluated.
The comparison shows whether the organization is mentioned spontaneously more often, whether it appears at higher positions, whether recommendations are better justified, which layers have improved, whether factual errors have disappeared, whether new sources are visibly being used, and whether competitors still win on the same arguments.
There is also a hard boundary in how conclusions are drawn. A change in measured score after interventions can be established. It cannot automatically be attributed entirely to those interventions. AI models change. Sources change. Competitors publish new information. A client may make their own adjustments outside the agreed plan. As I put it: those limitations don't belong in the fine print. They belong in the conclusion.
Frequently Asked Questions
What is the Agentability Framework?
The Agentability Framework is a structured measurement method developed by Jonkman Strategie that assesses how consistently an organization appears as a recommendation candidate across AI models like ChatGPT, Gemini, Claude, Perplexity and Grok. It runs 90 individual measurements per wave across five models, six frozen questions and three separate measurement days, scoring results against 25 checkpoints across five diagnostic layers.
What is the difference between AI recognition and AI recommendation candidacy?
Recognition means an AI model can describe an organization correctly when its name is mentioned directly. Recommendation candidacy means the model selects that organization spontaneously when a potential client asks a buying question without naming anyone. An organization can be well recognized but never appear in buying-context answers, which is the gap that actually affects client acquisition.
Why does AI visibility measurement use unmarked questions?
Unmarked questions omit the organization's name entirely, simulating how a real prospect searches for a provider. They reveal whether the organization appears as a spontaneous recommendation, not just whether it can be recognized when named. In one documented case by Jonkman Strategie, an organization was correctly recognized in marked questions but absent in 84 of 90 unmarked buying-question answers.
How long does an Agentability measurement cycle take?
A full cycle runs across two measurement waves separated by approximately 90 days. The first wave establishes the baseline across three measurement days. Interventions follow from the diagnosis. The second wave repeats the identical measurement conditions, with complete coding completed before any comparison to the baseline is made.
Can Agentability guarantee an increase in leads or revenue?
No. The Agentability Framework delivers a reproducible measurement of recommendation position, a layer-level diagnosis, a competitive benchmark and a documented comparison after interventions. It does not guarantee specific business outcomes, because AI models change independently and multiple external factors influence any shift in measured score.