Blog

ChatGPT

AI Visibility for SaaS and Plugins: When Does ChatGPT Recommend You?

Measure whether AI mentions, recommends, suggests, or invokes your SaaS product for real buyer jobs without collapsing different evidence into one score.

A SaaS company does not need another article explaining that ChatGPT can use apps.

It needs to know whether a buyer can describe a real job and encounter its product.

When a revenue operations lead asks for a tool to clean CRM duplicates, does the answer mention your platform? When a security-conscious enterprise buyer asks for an analytics product with SSO and data residency controls, are you recommended or excluded? When a connected capability is available, does ChatGPT merely name it, suggest connecting it, or actually use it?

Those are different commercial outcomes. They require different Questions, eligibility rules, evidence, and owners.

Short answer: Measure SaaS AI visibility from the buyer's job backward. Freeze a Question set, competitor cohort, channel, market, and account-state matrix. Then report ordinary brand mentions, explicit recommendations, suggested connections, actual invocations, and Plugin Directory discoverability as five separate evidence lanes. Do not turn them into one visibility score, and do not count a plan, permission, region, or connection failure as proof that the brand was ignored.

This guide provides a SaaS-specific Question framework, measurement contract, evidence schema, competitor workflow, and correction loop for doing that work.

The Commercial Question Is "Does AI Think Of Us For This Job?"

Most SaaS discovery happens before a buyer knows every vendor name. The buyer starts with a job:

  • Consolidate customer feedback from several channels.
  • Create sales collateral from an approved content library.
  • Find anomalous cloud spend before month-end.
  • Replace a spreadsheet-based onboarding process.
  • Analyze support tickets without exposing sensitive data.
  • Connect a CRM workflow to messaging and project systems.

A category-level mention tells you whether the brand is known. A job-level recommendation tells you whether the answer associates the brand with a particular need. A connection suggestion or invocation moves closer to product use. The measurement program should preserve that progression instead of labeling every event "AI visibility."

Brand teams should therefore ask:

  • Which buyer jobs produce our first appearance?
  • Which roles and company sizes change the shortlist?
  • Which integration, security, budget, or migration constraints exclude us?
  • Which competitors appear when we are absent?
  • Does the answer describe our fit accurately?
  • Is the evidence a text recommendation, a visible plugin suggestion, or a completed action?

These questions produce a backlog that marketing, product, partnerships, developer relations, security, and sales enablement can actually use.

Use Current OpenAI Terminology

OpenAI's terminology changed in 2026. Its current Plugins in ChatGPT and Codex documentation says that, as of July 9, 2026, the App Directory migrated to the Plugin Directory. A plugin can package skills, apps, and app templates. Apps remain the integrations that connect ChatGPT or Codex to external systems, data, and actions.

The distinction matters:

  • A company can be mentioned as a normal SaaS vendor without having an app.
  • An app can be included inside a plugin listing.
  • A visible listing does not prove that every user can install or invoke every included capability.
  • A connected app does not prove that ChatGPT will select it for every relevant request.

OpenAI's original apps in ChatGPT announcement says users can call an available app by name and that ChatGPT can also suggest an app when it is relevant to a conversation. That establishes contextual suggestion as an observable product behavior. It does not disclose a triggering formula, a guaranteed placement, or an impression-share report.

Use "Plugin Directory" for the current discovery surface and "app" for an integration that reaches external data or actions. Date historical "App Directory" evidence rather than silently renaming it.

Keep Five Evidence Lanes Separate

The following labels form an analyst's coding framework. They are not official OpenAI performance metrics.

Evidence LaneObservable EventBusiness QuestionDo Not Infer
Ordinary mentionThe answer names the brand or governed product aliasDoes AI associate us with this category or job?Endorsement, listing availability, or product use
Explicit recommendationThe answer affirmatively proposes the product for the stated needAre we entering the buyer's shortlist?A guaranteed ranking, connection, or conversion
Suggested connectionThe interface surfaces the relevant plugin or app and offers a selection or connection pathDoes the conversation create an opportunity to enable our capability?Successful authentication or invocation
Actual invocationThe app-backed capability runs and returns inspectable output or action evidenceCan an eligible user complete the job through the product?Universal availability or preference
Plugin Directory discoverabilityThe governed listing is found through a declared directory browse or search procedureCan an eligible user find the packaged capability intentionally?Contextual recommendation or invocation

An answer can satisfy more than one lane, but one event cannot substitute for another.

For example, "Tools such as Northstar and Beacon can analyze support tickets" is an ordinary mention for both brands. "Choose Northstar if you need EU data residency" is an explicit conditional recommendation. A Northstar plugin card with a Connect action is suggested-connection evidence. A completed analysis returned from Northstar is invocation evidence. Finding Northstar through the Plugin Directory is directory evidence.

Preserve the exact supporting sentence, visible interface state, or execution result. A binary code without the evidence is difficult to audit.

Separate Public Answers From Private Product State

Public answer monitoring and app execution are not the same experiment.

A public or clean answer baseline asks whether AI names and recommends a SaaS product from information available in the measured answer environment. A plugin test may depend on a logged-in account, eligible plan, supported surface, workspace policy, role, region, app enablement, OAuth connection, action controls, and permissions in the underlying SaaS system.

OpenAI's current Apps in ChatGPT documentation says that app capabilities can depend on plan, workspace, role, surface, and region. Its plugin guidance adds that an app-backed capability requires the underlying app to be enabled and that ChatGPT access does not override source-system permissions.

Create named state profiles before testing:

State ProfileAppropriate ObservationInappropriate Conclusion
Public answer baselineMention, recommendation, framing, displayed sourcesWhether a private app can connect
Eligible account, app not connectedVisible suggestion, selection path, connection offerWhether authentication will succeed
Eligible account, app connectedAvailability and actual invocation under declared permissionsAvailability to other roles or regions
Managed workspace, enabled roleRole-specific listing, connection, and invocationAvailability across the whole customer organization
Managed workspace, disabled rolePolicy or access block with reasonBrand absence or weak public visibility
Unsupported plan, surface, or regionEligibility resultProduct quality or recommendation failure

Do not pool these profiles. If ten public answers omit your brand and five managed-workspace attempts are blocked by policy, the report should show two outcomes. It should not call all fifteen "non-visible."

AEO Table is useful for the repeatable public answer lane: stable Questions, brands, competitors, Runs, answers, and citations. Account-specific suggestion and invocation testing belongs in a separate product QA lane, with screenshots, eligibility evidence, and app or plugin logs where the team is authorized to inspect them.

Build The Question Set From Buyer Jobs

Start with the buyer's desired outcome, not your homepage headline and not the word "plugin."

The AI search query set guide provides the general governance model. For vertical SaaS, annotate every Question across seven dimensions:

DimensionWhat To VaryExample
Job to be doneThe outcome the buyer needs"Find repeated product complaints across support tickets"
RoleThe person responsible for the job"for a customer success operations lead"
Company sizeOperating scale and complexity"for a 40-person startup" or "for a global enterprise"
IntegrationSystems the product must work with"using Salesforce, Slack, and Snowflake"
SecurityMaterial governance requirement"with SSO, audit logs, and EU data controls"
BudgetDeclared commercial constraint"under $500 per month" or "with usage-based pricing"
Switching constraintWhat must be replaced or migrated"without rebuilding our existing workflows"

Do not put every dimension into every Question. That creates unnatural prompts and makes diagnosis difficult. Use a stable base job, then add one or two constraints per variant.

Start With Unbranded Discovery Questions

Unbranded Questions reveal whether the product enters consideration before the buyer supplies its name:

  • "What tools can turn customer interviews into a prioritized product feedback report?"
  • "Which SaaS products help RevOps teams find duplicate and incomplete CRM records?"
  • "What software can generate sales presentations from an approved brand library?"
  • "Which support analytics tools are suitable for a small team with limited engineering resources?"
  • "What is a secure way to summarize internal research without copying files into a public chatbot?"

These should drive the headline result. They are more commercially meaningful than asking whether the answer knows your brand.

Add Role And Company-Size Variants

The same category can produce different shortlists for a solo operator and a regulated enterprise.

Examples:

  • "Best customer feedback tool for a product manager at an early-stage SaaS company."
  • "What analytics software can a global enterprise govern across several business units?"
  • "Which automation tool is realistic for a two-person operations team?"

Record the role and size facet as structured fields. Otherwise you will know that one Question changed but not which buyer constraint explains the pattern.

Add Integration And Security Constraints

Integration and security Questions test whether public evidence supports the decision:

  • "Which workflow automation tools connect to HubSpot and Microsoft Teams?"
  • "Which project tools offer SSO, audit logs, and role-based access control?"
  • "What AI writing tools publish clear data retention and training policies?"

Only include a requirement that materially changes vendor fit. Verify every resulting product claim against current public documentation before turning it into sales copy.

Add Budget And Switching Questions

Price and migration constraints expose a different part of the shortlist:

  • "Best reporting tool for a small agency with a $300 monthly software budget."
  • "What is a practical alternative to our spreadsheet onboarding tracker?"
  • "Which tools can replace a legacy survey platform without losing historical exports?"
  • "How can we move from manual competitor monitoring to a repeatable workflow?"

Do not assume that inclusion means the quoted price or migration description is correct. Code recommendation and claim accuracy separately.

Add Branded And Comparison Controls

Branded controls diagnose product understanding:

  • "What does Northstar Analytics do?"
  • "Does Northstar Analytics support SSO?"
  • "Northstar Analytics versus Beacon for a 100-person SaaS company."
  • "Alternatives to Northstar Analytics for teams using Microsoft 365."

Keep these out of the unbranded discovery rate. A brand supplied in the Question is expected to appear and would inflate a general visibility result.

Freeze A Competitor Cohort Before The Baseline

Choose a cohort that reflects how buyers solve the same job:

  • Direct category competitors.
  • A category leader that shapes buyer expectations.
  • A specialist for a key role or industry.
  • A low-cost or self-service alternative.
  • A substitute such as a spreadsheet, agency, internal build, or broad suite.

Define canonical names, product names, prior names, parent companies, and domains for every cohort member. Freeze the list before collecting results. Adding a competitor after reading the answers changes the denominator and makes the baseline difficult to reproduce.

Use the same cohort for each matched Question and Run. Code one binary appearance per brand per answer, even when the product name repeats. Then use the AI share-of-voice guide for explicit competitive mention and recommendation ratios. The broader competitor AI search tracking workflow explains how to inspect framing, source types, and content gaps.

Do not force every answer to have one winner. A useful response may recommend several products for different constraints. Preserve those conditions because they reveal the position each vendor occupies.

Write A Measurement Contract

The base unit for public answer monitoring is:

Question × answer surface × matched Run

For connection and invocation tests, add the state profile:

Question × answer surface × state profile × matched Run

The contract should declare:

  • Exact Question text and intent class.
  • Brand and competitor aliases.
  • Channel and product surface.
  • Market, language, and device.
  • Logged-in or public condition.
  • Plan and workspace type when relevant.
  • Role and plugin installation policy when relevant.
  • App enabled, connected, and authenticated state.
  • Underlying source-system permissions.
  • Search, retrieval, or tool mode when exposed.
  • Run window and retry policy.
  • Completion and eligibility rules.
  • Coding definitions and reviewer.

If the interface, directory structure, capability package, or availability policy changes, note a contract break. Do not quietly compare the new state with an old baseline as if the instrument were unchanged.

Establish A Baseline With Repeated Runs

One answer is an observation, not a trend.

Run the frozen core Question set repeatedly under matched conditions. Use enough repeated Runs to see whether an outcome recurs, while reporting the exact number rather than calling an arbitrary count statistically conclusive. Keep campaign and exploratory Questions outside the core baseline.

For each period, report:

  • Attempted units.
  • Completed eligible answers.
  • Failed, refused, blocked, and ineligible units.
  • Mention and recommendation outcomes.
  • Suggestion, connection, and invocation eligibility.
  • Directory checks by declared profile.
  • Changes to product, content, listing, or account state.

The AI search volatility guide explains why a single appearance or disappearance should not be treated as durable movement. Repeat the matched baseline before and after a meaningful change, and keep the raw evidence for both windows.

Calculate Separate Rates With Separate Denominators

Use a small dashboard of explicit rates.

MetricNumeratorDenominator
Public mention coverageCompleted eligible public answers naming the brandCompleted eligible public answers
Explicit recommendation coverageCompleted eligible public answers affirmatively recommending the brandCompleted eligible public answers
Competitive recommendation shareBinary target recommendationsBinary recommendations for every brand in the frozen cohort
Suggested-connection observation rateEligible attempts showing a governed suggestion or connection pathEligible attempts in the same state profile
Invocation success rateEligible invocation attempts returning the declared success evidenceEligible invocation attempts
Directory discovery rateDeclared directory procedures that find the governed listingEligible directory procedures under the same profile

Show N/A when no eligible denominator exists. Do not convert it to zero.

Never average these six rates into one "SaaS AI visibility score." A company could have strong public recommendations but no app, excellent directory discovery but weak public category association, or reliable invocation limited to one permitted workspace role. The separate lanes tell the team what to fix.

Preserve Evidence That Another Reviewer Can Audit

For every observation, store enough context to reproduce the classification:

FieldEvidence To Preserve
Observation IDStable identifier linking Question, profile, and Run
Buyer intentJob, role, size, integration, security, budget, and switching facets
Exact QuestionFull submitted text without post-Run editing
Surface and stateProduct surface, market, language, device, plan, workspace, and role
EligibilityEligible, failed, refused, policy-blocked, unavailable, or other declared reason
Answer evidenceRaw answer text and displayed sources
Mention codeBrand alias matched and supporting sentence
Recommendation codeRecommended, listed, compared, caveated, or absent, with excerpt
Suggestion evidenceVisible plugin or app suggestion and connection path
Invocation evidenceSelected capability, confirmation state, output or action result, and error
Directory evidenceSearch or browse procedure, query, filters, position if visible, and screenshot
Competitor evidenceBrands, products, framing, citations, and governed domains
Review dataTimestamp, coding version, reviewer, and adjudication note

Redact credentials, tokens, private customer data, and sensitive records. Evidence quality does not require exposing secrets.

Diagnose The Competitor Gap

After the baseline, group absence and competitor wins by buyer intent.

Ask:

  • Which competitor owns the broad category Questions?
  • Which vendor appears only after an enterprise security constraint?
  • Which product wins integration-specific Questions?
  • Is a low-cost competitor recommended only when budget is explicit?
  • Does a broad suite replace specialists for larger companies?
  • Which competitors receive suggestions or invocations in eligible states?
  • Which sources support the competitor's claims?
  • Is our product absent, incorrectly described, or correctly excluded?

The last question prevents wasted work. If your product genuinely lacks a required integration or residency option, the answer may be accurately excluding it. Marketing should not publish vague copy to imply otherwise. Product can decide whether the gap belongs on the roadmap.

Separate at least four diagnoses:

DiagnosisObservable PatternLikely Owner
Entity gapThe product is absent or confused even in relevant branded controlsBrand, SEO, and content operations
Evidence gapThe answer cannot verify integrations, security, pricing, or migration claimsProduct marketing, docs, security, and legal
Positioning gapCompetitors are consistently assigned a clear job while your product is described genericallyProduct marketing and category strategy
Product or access gapListing, connection, permission, or invocation fails under an eligible profileProduct, engineering, partnerships, and workspace admin

Turn Findings Into Content And Product Work

Build the backlog around observed buyer constraints.

Strengthen Job Pages

Create pages that answer one real job for one audience. State the workflow, inputs, outputs, limitations, setup, proof, and next step.

Publish Integration Evidence

Explain what connects, data direction, required plans or permissions, setup, supported actions, limitations, and update status. A logo wall cannot support a detailed recommendation.

Make Security Claims Reviewable

Maintain current public information for access controls, audit logs, retention, data use, residency, certifications, subprocessors, and deletion. Use precise scope and dates.

Clarify Pricing And Buyer Fit

Explain packaging units, limits, trial conditions, add-ons, and who each plan fits. If pricing requires a quote, state the deciding factors.

Build Switching And Comparison Resources

Document migration inputs, export formats, implementation effort, and limitations. Comparison pages should use governed facts and explain fit by constraint.

Review The Plugin Listing And Connection Path

Verify that the listing describes the workflow, included app, setup, permissions, terms, and privacy policy. Test connection with an eligible low-risk account before invocation.

Validate The Product Experience

If invocation fails, preserve the request, capability, confirmation state, error, available app logs, and permission result. Content cannot repair broken authentication or an unavailable action.

After a change, rerun only the matched Questions and state profiles needed to test the hypothesis, then repeat the full core baseline on its normal cadence. An observed improvement after publication is evidence of change, not proof that the edit caused an undisclosed recommendation system to prefer the product.

Use A Four-Week Operating Cycle

Week 1: Contract And Baseline

Freeze Questions, competitors, profiles, coding rules, and denominators. Run the matched baseline.

Week 2: Gap Review

Group results by buyer constraint. Separate accurate exclusion from evidence, positioning, and product failures.

Week 3: One Governed Change Cluster

Update one evidence cluster or connection path. Record URLs, product versions, owners, and release times.

Week 4: Matched Validation

Repeat the same Questions and profiles. Report completed, failed, blocked, and ineligible units without changing the cohort.

Continue the stable core on a recurring cadence. Add new buyer language to an exploration set first, then promote it into the core only through a documented version change.

Common Measurement Mistakes

Searching Only For The Brand Name

Branded Questions test recognition, not unbranded discovery.

Treating A Mention As A Recommendation

A neutral list, comparison, or warning is not an endorsement. Preserve the supporting sentence.

Calling A Suggestion An Invocation

A plugin card or Connect button is an opportunity. Invocation requires execution evidence.

Counting Permission Blocks As Brand Absence

Record plan, region, role, policy, connection, or permission blocks as eligibility outcomes.

Mixing Public And Connected Runs

Keep public, unconnected eligible, and connected eligible profiles separate.

Changing Competitors After Seeing Results

A revised cohort changes the denominator. Version it instead of rewriting history.

Optimizing For A Hidden Trigger

OpenAI documents contextual suggestions, not a trigger formula or guarantee. Improve evidence, listing clarity, permissions, and product usefulness; then measure.

Publishing One Combined Score

A blended score conceals whether the gap is awareness, discovery, connection, or execution.

The Bottom Line

SaaS AI visibility is a set of buyer-centered questions:

  • Does the product appear for the jobs it can genuinely complete?
  • Is it recommended for the right roles, company sizes, and constraints?
  • Are competitors framed more clearly or supported by better evidence?
  • Can eligible users find and connect the packaged capability?
  • Does the app actually complete the declared job under the required permissions?

Build a stable Question set, freeze the cohort, and separate public answers from private account state. Preserve each outcome with its own denominator, then turn observed gaps into accurate content, positioning, documentation, and product work.

Track SaaS mentions and recommendations with AEO Table and establish the public answer baseline before adding account-specific plugin QA.

FAQ

How can a SaaS company measure whether ChatGPT recommends its product?

Build a stable set of unbranded buyer Questions, run them under declared channel and account conditions, and preserve the answers. Code ordinary mentions, explicit recommendations, suggested connections, actual invocations, and directory discovery separately, with visible eligible denominators.

Is a ChatGPT mention the same as a plugin invocation?

No. A text answer can name or recommend a SaaS brand without surfacing its plugin or invoking an app. Invocation requires separate evidence that the relevant capability actually ran under an eligible, connected, and permitted account state.

What affects whether a ChatGPT plugin is available to a user?

OpenAI says availability can depend on plan, workspace settings, role, supported surface, region, included capabilities, app enablement, authentication, and source-system permissions. Record these conditions instead of treating an unavailable plugin as zero brand visibility.

Can AEO work guarantee that ChatGPT will suggest or invoke a SaaS app?

No. OpenAI documents contextual app suggestions but does not publish a trigger formula or guarantee a suggestion or invocation. AEO teams can improve public product evidence and measure observed outcomes, while product teams separately validate listing, connection, permissions, and app execution.