Research
One public URL in. A classified list of other URLs out. Then scrape those
pages for who they sell to and how they go to market. Do not turn a research
hit into a Salesforce account, a competitor_ids bind, or call language.
This page is a tactics collection for agents. It is not a generated GTM guide and not a canonical competitor profile. For installed-base harvest from Certificate Transparency, use Domain scanning. For a writeable competitor record, use the sales competitor workflow with a name plus G2 and Capterra links.
Tools
Prefer tools an agent can call. Cloudflare Browser Run and Firecrawl are the default stack. Add a specialist only when the job needs it.
| Tool | Agent surface | Use for |
|---|---|---|
| Cloudflare Browser Run (formerly Browser Rendering) | REST Quick Actions, Workers env.BROWSER.quickAction(), Playwright MCP, Stagehand | Render a page after JS. Markdown, structured JSON, link lists, screenshots, shallow crawl. |
| Firecrawl | REST /v2/search, /v2/map, /v2/scrape, /v2/crawl; MCP | Web search, sitemap-speed URL maps, scrape to markdown/JSON, site crawl. |
| Exa | REST /search, /findSimilar, /contents; Company search; Websets | Pages like a seed URL. Company lists. Optional Similarweb via Exa Connect. |
| crt.name | GET /v1/search?apex= | Customer hostnames of a multi-tenant SaaS. Then lookup. See Domain scanning. |
| Cloudflare DNS (1.1.1.1 DoH) | https://cloudflare-dns.com/dns-query | Does {label}.com exist after a CT slug or a guessed vendor apex? |
| Wayback CDX | https://web.archive.org/cdx/search/cdx?url= | Old customers, pricing, and positioning pages that the live site dropped. |
Cloudflare Quick Action base:
POST https://api.cloudflare.com/client/v4/accounts/<accountId>/browser-rendering/{markdown|json|links|screenshot|crawl}
Authorization: Bearer <token with Browser Rendering - Edit>
Firecrawl v2 base:
POST https://api.firecrawl.dev/v2/{search|map|scrape|crawl}
Authorization: Bearer <fc key>
Exa base:
POST https://api.exa.ai/{search|findSimilar|contents}
x-api-key: <exa key>
Default loop: Browser Run to read a known page. Firecrawl to discover more URLs. Exa when the seed is one URL and the job is "more like this." crt.name when the seed is a SaaS apex and the job is "who is a tenant."
Context
Use this page when the question is one of:
- Who already uses this vendor? Give me their websites.
- Who else sells in this category? Give me vendor websites.
- This page is a good example. Find pages like it.
- What is this vendor actually doing in market (offers, proof, channels, hiring, content)?
Do not use this page to invent a competitor for an audience YAML, to claim a plant runs a named ERP, or to scrape behind a login.
Axioms
- Public pages only. No customer login, no stolen session, no paywall bypass.
- A URL is a research candidate. Identity, ICP, and system-of-record still have to be confirmed.
- Prefer one tool per step. Do not crawl a whole internet when
/mapor/searchis enough. - Rate-limit. Honor
robots.txtwhen the tool does not. - Write findings as hypotheses with the source URL. Never as account facts.
1. Find customer URLs
Goal: websites of companies that appear to use a named vendor. Not proof they use it today.
Work the methods in order. Stop when the list is large enough to qualify.
A. Vendor proof pages (fastest)
The vendor already published names. Map the site, keep proof URLs, scrape names, then resolve each name to a domain.
-
Firecrawl map with a topic filter:
POST /v2/map{ "url": "https://{vendor-apex}", "search": "customers case study story logo" }Keep paths that look like
/customers,/case-stud,/stories,/proof,/resources/,/partners. Drop/blogunless the post is a named customer. -
Cloudflare
/jsonon each kept page. Schema:{"page_type": "case_study | logo_wall | partner_dir | other","customer_names": ["..."],"industries_mentioned": ["..."],"quoted_outcomes": ["..."]} -
For each
customer_name, Firecrawl/search:{ "query": "\"{customer_name}\" official website", "limit": 5 }Confirm the domain on the company's own site (title, legal name, footer). Cloudflare DNS if you only have a slug.
B. Tenant hostnames (SaaS installed base)
If the product issues {tenant}.{vendor-apex} certificates, run
Domain scanning. That is the customer-URL method
for multi-tenant vendors. Do not duplicate the scanner here.
C. "Uses X" on the open web
Firecrawl /search (optionally with scrapeOptions for markdown):
"{vendor product}" (customer OR "case study" OR "powered by" OR "we use")
"{vendor product}" "injection molding"
site:g2.com/products/{slug}/reviews
Pull company names from reviews and posts. Resolve to domains as in A.3. G2 and Capterra are name sources. They are not the company's website.
D. Jobs and partners
Search {vendor} "implementation" OR careers OR "we are hiring".
Implementation partners often list their customers; treat those as a
second-hand list and resolve each name to a domain.
Done when: each row has customer_name, company_domain,
evidence_url, method (vendor_proof | ct_tenant | web_mention |
partner), confidence (guessed | page_confirmed).
2. Find vendor URLs
Goal: other product/company websites in the same category as a seed vendor. Example seed we already own: Global Shop.
A. Category directories
Firecrawl /search:
"{category} software" (ERP OR MES OR "shop floor") {vertical}
"alternatives to {vendor}"
site:g2.com/products {category}
site:capterra.com {category}
Take product homepages, not the directory URL. A G2 slug is a pointer. The vendor URL is the company's own apex.
B. Similar vendor sites
Take one known vendor homepage. Run §3
with page_type: vendor_home. Keep results that are other vendors, not
their customers.
C. Integration and marketplace graphs
Map a platform the seed integrates with (/integrations, /partners,
AppExchange-style directories). Cloudflare /links or Firecrawl /map
with search: "erp manufacturing". Each listed product is a vendor
candidate.
D. Our catalog, then widen
Start from Competitors and
system_of_record in variables.yaml. Research
may propose a new name. It does not create C_xxxxx or add
competitor_ids. Canonical profiles still need the competitor workflow
(name + G2 + Capterra).
Done when: each row has vendor_name, vendor_domain, evidence_url,
why_same_category (one sentence from the page), already_in_catalog
(yes/no).
3. Similar pages from one URL
Goal: given seed_url, return other URLs of the same kind (another
molder's site, another ERP vendor, another case study, another pricing
page).
Fingerprint the seed first. Then search. Do not search the raw URL string and hope.
Step 1 — Fingerprint
Cloudflare /json or Firecrawl scrape with JSON format:
{
"page_type": "vendor_home | product | pricing | case_study | customer_home | careers | other",
"entity_name": "",
"category": "",
"process_or_industry": "",
"geo": "",
"distinctive_phrases": [""],
"named_systems": [""]
}
customer_home is a plant or operating company. vendor_home is
software. Do not mix them in one similar-page run.
Step 2 — Retrieve (pick one)
Exa, when the job is literally "URLs like this one." /findSimilar
is deprecated in favor of /search with a description of the source, but
it still exists for a raw URL:
POST https://api.exa.ai/findSimilar
{ "url": "{seed_url}", "numResults": 20, "excludeSourceDomain": true }
Preferred Exa shape: /search with a query built from the fingerprint,
not the URL.
Firecrawl, default. Search with the fingerprint, not the URL:
POST /v2/search
{
"query": "{category} {process_or_industry} {page_type} {geo} \"{distinctive_phrase}\"",
"limit": 20
}
If the seed is a vendor, add "alternatives" OR software OR ERP. If the
seed is a customer plant, add the process (injection molding, CNC)
and drop software words.
Cloudflare-only fallback. /markdown the seed, /links for
outbound neighbors, then Firecrawl search on the best phrase in the
markdown. Browser Run has no native similar-page index.
Step 3 — Filter
Scrape each result enough to re-run the fingerprint. Keep when
page_type matches and category or process_or_industry matches.
Drop the seed domain, socials, and directory pages (G2, Capterra, LinkedIn)
unless you are collecting vendor names.
Done when: a list of {url, page_type, entity_name, similarity_note}.
similarity_note is one line: what matched (process, category, offer).
4. Read competitor tactics
Goal: from a vendor URL, extract observable go-to-market moves. Not a battlecard. Not proof we beat them.
Map once, scrape the small set below.
Map
POST /v2/map
{ "url": "https://{vendor-apex}", "search": "pricing customers blog integrations partners careers changelog webinar" }
Or Cloudflare /crawl if you already live on Workers and the site is
small.
Read this set
| Page | What to extract |
|---|---|
| Home / product | Category claim, ICP language, named modules |
| Pricing | Published packages, "contact us", what's free vs paid |
| Customers / case studies | Named logos, verticals, outcomes they chose to show |
| Blog / resources / webinars | Topics they rank and teach. Frequency. |
| Integrations / partners | What they attach to. Channel motion. |
| Careers | Roles they are hiring (AE, implementation, CS). Investment, not gossip. |
| Changelog / docs | What they just shipped. |
| Legal / security | Enterprise readiness signals |
Cloudflare /json or Firecrawl JSON mode, one schema for the vendor:
{
"positioning_one_liner": "",
"named_icp": [""],
"offers": [""],
"proof_customers": [""],
"channels_visible": ["content | partner | webinar | ads | unknown"],
"integrations": [""],
"hiring_motions": [""],
"pricing_shape": "public | sales_led | freemium | unknown",
"source_urls": [""]
}
Optional: Firecrawl change tracking or a scheduled /scrape on pricing
and home. Optional: Wayback CDX on the same paths for what they used to
claim.
Ads (Meta Ad Library, Google Ads Transparency) are public but clumsy for agents. Use them only when you already have the advertiser name and need creative examples. Do not block the rest of the read on ads.
Done when: a one-page tactics note with source URLs. A canonical profile still needs the sales competitor workflow: a name plus G2 and Capterra review links. See Competitors. This note is not that profile.
Agent runbook
One pass. Do not interleave jobs.
- Name the job:
customers|vendors|similar|tactics. - Name the seed: vendor apex, or one
seed_url. - Discover with Firecrawl
/searchor/map, Exa/findSimilar, or crt.name. Not all four. - Read with Cloudflare
/markdown+/json, or Firecrawl/scrape. - Classify every URL. Drop vendor infrastructure and directories.
- Write a table. Stop. Qualification and CRM are other pages.
If /map returns almost nothing, the site has a weak sitemap. Switch to
Cloudflare /links on the homepage, then /crawl, or Firecrawl /crawl.
If JSON extract hallucinates customer names, keep only names that appear as visible text in the markdown. Re-scrape; do not trust the schema output alone.
What this is not
- Permission to say on a call that an account uses Global Shop, or any other system.
- A substitute for ZoomInfo identity and phones (Sales stack).
- A substitute for Domain scanning when the artifact is a tenant hostname.
- Authority to add
competitor_idsor mintC_xxxxx.
Related knowledge
- Comparison research — G2-class review and alternatives directories
- GEO — engines cite these hosts; scrape what they ranked
- AEO — coding agents fetch markdown and registries
- ERP signal research — 50–500 ERP need / want / window sources
- Capital signal research — PE firm lists and portfolio scrapes
- Expo signal research — exhibitor directories and attendance lists
- CEO scene videos — one public artifact into a Voss-structured video to a named P1
- Project bidding — SAM, BidNet, cooperatives, issuing portals; find vs bid
- Domain scanning — CT harvest of vendor tenants
- Competitors
- GTM guide structure — section 1.3 market landscape is a request for sourced content, not a license to invent
- Sales stack
Change log
- 2026-08-31 — First tactics page: agentic tool table, customer URLs, vendor URLs, similar-page fingerprint, competitor tactic extract. Defaults to Cloudflare Browser Run and Firecrawl; Exa and crt.name as specialists.