Skip to main content

Research

One public URL in. A classified list of other URLs out. Then scrape those pages for who they sell to and how they go to market. Do not turn a research hit into a Salesforce account, a competitor_ids bind, or call language.

This page is a tactics collection for agents. It is not a generated GTM guide and not a canonical competitor profile. For installed-base harvest from Certificate Transparency, use Domain scanning. For a writeable competitor record, use the sales competitor workflow with a name plus G2 and Capterra links.

Tools​

Prefer tools an agent can call. Cloudflare Browser Run and Firecrawl are the default stack. Add a specialist only when the job needs it.

ToolAgent surfaceUse for
Cloudflare Browser Run (formerly Browser Rendering)REST Quick Actions, Workers env.BROWSER.quickAction(), Playwright MCP, StagehandRender a page after JS. Markdown, structured JSON, link lists, screenshots, shallow crawl.
FirecrawlREST /v2/search, /v2/map, /v2/scrape, /v2/crawl; MCPWeb search, sitemap-speed URL maps, scrape to markdown/JSON, site crawl.
ExaREST /search, /findSimilar, /contents; Company search; WebsetsPages like a seed URL. Company lists. Optional Similarweb via Exa Connect.
crt.nameGET /v1/search?apex=Customer hostnames of a multi-tenant SaaS. Then lookup. See Domain scanning.
Cloudflare DNS (1.1.1.1 DoH)https://cloudflare-dns.com/dns-queryDoes {label}.com exist after a CT slug or a guessed vendor apex?
Wayback CDXhttps://web.archive.org/cdx/search/cdx?url=Old customers, pricing, and positioning pages that the live site dropped.

Cloudflare Quick Action base:

POST https://api.cloudflare.com/client/v4/accounts/<accountId>/browser-rendering/{markdown|json|links|screenshot|crawl}
Authorization: Bearer <token with Browser Rendering - Edit>

Firecrawl v2 base:

POST https://api.firecrawl.dev/v2/{search|map|scrape|crawl}
Authorization: Bearer <fc key>

Exa base:

POST https://api.exa.ai/{search|findSimilar|contents}
x-api-key: <exa key>

Default loop: Browser Run to read a known page. Firecrawl to discover more URLs. Exa when the seed is one URL and the job is "more like this." crt.name when the seed is a SaaS apex and the job is "who is a tenant."

Context​

Use this page when the question is one of:

  1. Who already uses this vendor? Give me their websites.
  2. Who else sells in this category? Give me vendor websites.
  3. This page is a good example. Find pages like it.
  4. What is this vendor actually doing in market (offers, proof, channels, hiring, content)?

Do not use this page to invent a competitor for an audience YAML, to claim a plant runs a named ERP, or to scrape behind a login.

Axioms​

  • Public pages only. No customer login, no stolen session, no paywall bypass.
  • A URL is a research candidate. Identity, ICP, and system-of-record still have to be confirmed.
  • Prefer one tool per step. Do not crawl a whole internet when /map or /search is enough.
  • Rate-limit. Honor robots.txt when the tool does not.
  • Write findings as hypotheses with the source URL. Never as account facts.

1. Find customer URLs​

Goal: websites of companies that appear to use a named vendor. Not proof they use it today.

Work the methods in order. Stop when the list is large enough to qualify.

A. Vendor proof pages (fastest)​

The vendor already published names. Map the site, keep proof URLs, scrape names, then resolve each name to a domain.

  1. Firecrawl map with a topic filter:

    POST /v2/map
    { "url": "https://{vendor-apex}", "search": "customers case study story logo" }

    Keep paths that look like /customers, /case-stud, /stories, /proof, /resources/, /partners. Drop /blog unless the post is a named customer.

  2. Cloudflare /json on each kept page. Schema:

    {
    "page_type": "case_study | logo_wall | partner_dir | other",
    "customer_names": ["..."],
    "industries_mentioned": ["..."],
    "quoted_outcomes": ["..."]
    }
  3. For each customer_name, Firecrawl /search:

    { "query": "\"{customer_name}\" official website", "limit": 5 }

    Confirm the domain on the company's own site (title, legal name, footer). Cloudflare DNS if you only have a slug.

B. Tenant hostnames (SaaS installed base)​

If the product issues {tenant}.{vendor-apex} certificates, run Domain scanning. That is the customer-URL method for multi-tenant vendors. Do not duplicate the scanner here.

C. "Uses X" on the open web​

Firecrawl /search (optionally with scrapeOptions for markdown):

"{vendor product}" (customer OR "case study" OR "powered by" OR "we use")
"{vendor product}" "injection molding"
site:g2.com/products/{slug}/reviews

Pull company names from reviews and posts. Resolve to domains as in A.3. G2 and Capterra are name sources. They are not the company's website.

D. Jobs and partners​

Search {vendor} "implementation" OR careers OR "we are hiring". Implementation partners often list their customers; treat those as a second-hand list and resolve each name to a domain.

Done when: each row has customer_name, company_domain, evidence_url, method (vendor_proof | ct_tenant | web_mention | partner), confidence (guessed | page_confirmed).

2. Find vendor URLs​

Goal: other product/company websites in the same category as a seed vendor. Example seed we already own: Global Shop.

A. Category directories​

Firecrawl /search:

"{category} software" (ERP OR MES OR "shop floor") {vertical}
"alternatives to {vendor}"
site:g2.com/products {category}
site:capterra.com {category}

Take product homepages, not the directory URL. A G2 slug is a pointer. The vendor URL is the company's own apex.

B. Similar vendor sites​

Take one known vendor homepage. Run §3 with page_type: vendor_home. Keep results that are other vendors, not their customers.

C. Integration and marketplace graphs​

Map a platform the seed integrates with (/integrations, /partners, AppExchange-style directories). Cloudflare /links or Firecrawl /map with search: "erp manufacturing". Each listed product is a vendor candidate.

D. Our catalog, then widen​

Start from Competitors and system_of_record in variables.yaml. Research may propose a new name. It does not create C_xxxxx or add competitor_ids. Canonical profiles still need the competitor workflow (name + G2 + Capterra).

Done when: each row has vendor_name, vendor_domain, evidence_url, why_same_category (one sentence from the page), already_in_catalog (yes/no).

3. Similar pages from one URL​

Goal: given seed_url, return other URLs of the same kind (another molder's site, another ERP vendor, another case study, another pricing page).

Fingerprint the seed first. Then search. Do not search the raw URL string and hope.

Step 1 — Fingerprint​

Cloudflare /json or Firecrawl scrape with JSON format:

{
"page_type": "vendor_home | product | pricing | case_study | customer_home | careers | other",
"entity_name": "",
"category": "",
"process_or_industry": "",
"geo": "",
"distinctive_phrases": [""],
"named_systems": [""]
}

customer_home is a plant or operating company. vendor_home is software. Do not mix them in one similar-page run.

Step 2 — Retrieve (pick one)​

Exa, when the job is literally "URLs like this one." /findSimilar is deprecated in favor of /search with a description of the source, but it still exists for a raw URL:

POST https://api.exa.ai/findSimilar
{ "url": "{seed_url}", "numResults": 20, "excludeSourceDomain": true }

Preferred Exa shape: /search with a query built from the fingerprint, not the URL.

Firecrawl, default. Search with the fingerprint, not the URL:

POST /v2/search
{
"query": "{category} {process_or_industry} {page_type} {geo} \"{distinctive_phrase}\"",
"limit": 20
}

If the seed is a vendor, add "alternatives" OR software OR ERP. If the seed is a customer plant, add the process (injection molding, CNC) and drop software words.

Cloudflare-only fallback. /markdown the seed, /links for outbound neighbors, then Firecrawl search on the best phrase in the markdown. Browser Run has no native similar-page index.

Step 3 — Filter​

Scrape each result enough to re-run the fingerprint. Keep when page_type matches and category or process_or_industry matches. Drop the seed domain, socials, and directory pages (G2, Capterra, LinkedIn) unless you are collecting vendor names.

Done when: a list of {url, page_type, entity_name, similarity_note}. similarity_note is one line: what matched (process, category, offer).

4. Read competitor tactics​

Goal: from a vendor URL, extract observable go-to-market moves. Not a battlecard. Not proof we beat them.

Map once, scrape the small set below.

Map​

POST /v2/map
{ "url": "https://{vendor-apex}", "search": "pricing customers blog integrations partners careers changelog webinar" }

Or Cloudflare /crawl if you already live on Workers and the site is small.

Read this set​

PageWhat to extract
Home / productCategory claim, ICP language, named modules
PricingPublished packages, "contact us", what's free vs paid
Customers / case studiesNamed logos, verticals, outcomes they chose to show
Blog / resources / webinarsTopics they rank and teach. Frequency.
Integrations / partnersWhat they attach to. Channel motion.
CareersRoles they are hiring (AE, implementation, CS). Investment, not gossip.
Changelog / docsWhat they just shipped.
Legal / securityEnterprise readiness signals

Cloudflare /json or Firecrawl JSON mode, one schema for the vendor:

{
"positioning_one_liner": "",
"named_icp": [""],
"offers": [""],
"proof_customers": [""],
"channels_visible": ["content | partner | webinar | ads | unknown"],
"integrations": [""],
"hiring_motions": [""],
"pricing_shape": "public | sales_led | freemium | unknown",
"source_urls": [""]
}

Optional: Firecrawl change tracking or a scheduled /scrape on pricing and home. Optional: Wayback CDX on the same paths for what they used to claim.

Ads (Meta Ad Library, Google Ads Transparency) are public but clumsy for agents. Use them only when you already have the advertiser name and need creative examples. Do not block the rest of the read on ads.

Done when: a one-page tactics note with source URLs. A canonical profile still needs the sales competitor workflow: a name plus G2 and Capterra review links. See Competitors. This note is not that profile.

Agent runbook​

One pass. Do not interleave jobs.

  1. Name the job: customers | vendors | similar | tactics.
  2. Name the seed: vendor apex, or one seed_url.
  3. Discover with Firecrawl /search or /map, Exa /findSimilar, or crt.name. Not all four.
  4. Read with Cloudflare /markdown + /json, or Firecrawl /scrape.
  5. Classify every URL. Drop vendor infrastructure and directories.
  6. Write a table. Stop. Qualification and CRM are other pages.

If /map returns almost nothing, the site has a weak sitemap. Switch to Cloudflare /links on the homepage, then /crawl, or Firecrawl /crawl.

If JSON extract hallucinates customer names, keep only names that appear as visible text in the markdown. Re-scrape; do not trust the schema output alone.

What this is not​

  • Permission to say on a call that an account uses Global Shop, or any other system.
  • A substitute for ZoomInfo identity and phones (Sales stack).
  • A substitute for Domain scanning when the artifact is a tenant hostname.
  • Authority to add competitor_ids or mint C_xxxxx.

Change log​

  • 2026-08-31 — First tactics page: agentic tool table, customer URLs, vendor URLs, similar-page fingerprint, competitor tactic extract. Defaults to Cloudflare Browser Run and Firecrawl; Exa and crt.name as specialists.