Company & Startup Intelligence Tools β Sourcing, Funding & Research Data
Structured, API-grade data on private companies, startup ecosystems, VC portfolios and founder networks β built for sourcing analysts, recruiters, sales teams and BD operators who need fresher signal than Crunchbase can deliver.
What’s in this category
The Company & Startup Intelligence toolkit covers the full private-market data stack: accelerator alumni databases (Y Combinator, Techstars, 500 Global), bootstrapper and indie-product directories (Indie Hackers, Product Hunt), VC portfolio scrapers (Founders Fund, Lightspeed, Index Ventures, Greylock, Bessemer, a16z, Sequoia), live funding feeds from SEC EDGAR Form D filings and TechCrunch, and company enrichment utilities that turn a bare company name or domain into a complete profile with employee count, tech stack, social presence and contact data. Adjacent vertical signals β Shopify merchant detection, GitHub trending repos, healthcare provider directories β round out the lineup for operators working specific industry slices.
Who uses it: VC associates running deal sourcing, executive recruiters building target maps, sales reps prospecting Series A startups, BD teams pursuing partnership pipelines, journalists tracking funding announcements, M&A scouts mapping consolidation targets, and growth marketers seeded against accelerator batches. If your job involves answering “which companies just got funded, who founded them, and how do I reach them?” β this category is the spine of your workflow. Everything here returns structured JSON or CSV, runs on pay-per-result pricing, and integrates cleanly with downstream tools like Snowflake, BigQuery, HubSpot, Salesforce, Clay, Hightouch and Apollo.
Common use cases
- Source target companies for Series A outreach β pull the latest YC batch combined with recent SEC Form D filings, filter by stage and vertical, hand the deduplicated list to your SDRs with verified domains attached.
- Build a Founders Fund / Lightspeed / Greylock portfolio coverage list β combine multiple VC scrapers into a single deduplicated coverage universe for your fund’s competitive intelligence team or your sales territory mapping.
- Track YC W26 batch for sales prospecting β schedule the YC Companies Directory weekly, diff against last week’s run, alert your team via Slack on every new batch addition with website and one-liner attached.
- Identify newly-funded startups in a vertical β combine the Startup Funding Tracker with industry keyword filters to surface fintech, healthtech or AI rounds within hours of the SEC Form D filing, well before TechCrunch picks it up.
- Enrich a CRM with founder names, domains and emails β feed a list of company names into Company Enrichment, get back fully-hydrated profiles ready to push to HubSpot or Salesforce via reverse-ETL.
- Map a competitor’s entire customer base β use Shopify App Merchant Finder to reverse-lookup every store running a given app (Klaviyo, Recharge, Gorgias), then enrich the merchant list with revenue estimates.
- Catch acquisition signals across a VC portfolio β schedule every portfolio scraper daily, diff the “status” field, alert on any company that flips from “Active” to “Acquired” or “IPO” before the news embargo lifts.
- Build a recruiter target map β combine Greylock + Bessemer + a16z portfolios with GitHub Trending to find ML-heavy startups hiring senior engineers, then enrich with founder emails for direct outreach.
Featured tools
All actors are public on Apify, run on pay-per-result pricing, and ship with documented input schemas plus example outputs. The five VC portfolio scrapers tagged [NEW] just flipped public today β Founders Fund, Lightspeed, Index Ventures, Greylock and Bessemer are the freshest additions to the lineup and represent the first structured feeds of these portfolios available on Apify.
| Actor | Type | What it does |
|---|---|---|
| Founders Fund Portfolio Scraper [NEW] | VC portfolio | Peter Thiel’s VC: SpaceX, Palantir, Anduril, Stripe, OpenAI, Neuralink, Ramp and 200+ companies with stage, sector, status. |
| Lightspeed Portfolio Scraper [NEW] | VC portfolio | 600+ Lightspeed Venture Partners companies with stage invested, partner names, founders, status, and industry. |
| Index Ventures Portfolio Scraper [NEW] | VC portfolio | The only structured Index Ventures portfolio feed β ~386 cross-Atlantic companies with sector and stage attribution. |
| Greylock Portfolio Scraper [NEW] | VC portfolio | Complete Greylock Partners portfolio (Reid Hoffman, David Sze, Saam Motamedi, Sarah Guo) with sector and partner attribution. |
| Bessemer Venture Partners Portfolio Scraper [NEW] | VC portfolio | 500+ Bessemer companies with their unique Roadmap sector taxonomy, partner names, and current investment status. |
| a16z Portfolio Scraper | VC portfolio | Canonical scraper for the full Andreessen Horowitz portfolio: 800+ companies with sector, stage, year invested, and founder data. |
| Sequoia Capital Portfolio Scraper | VC portfolio | The only structured Sequoia + Peak XV portfolio feed β ~710 companies across US, India, and SEA funds. |
| YC Companies Directory | Accelerator | Every Y Combinator batch from S05 to today β id, name, slug, website, one-liner, batch tag, current status, team size. |
| Startup Funding Tracker | Funding feed | Series A/B funding rounds aggregated from SEC EDGAR Form D filings, TechCrunch and YC β a real Crunchbase alternative. |
| Company Enrichment Tool | Enrichment | Enrich company names with domains, emails, social profiles, employee count, industry and tech stack β B2B enrichment for CRM hydration. |
| Company Data Aggregator | Enrichment | Bulk company profile lookup. Aggregates WHOIS, DNS, GitHub org, SSL certs, tech stack headers β zero auth, Crunchbase API alternative. |
| Company Logo API | Utility | Drop-in Clearbit Logo API replacement (sunset Dec 2025). Bulk-fetch company logos by domain via a 5-source fallback chain. |
| Shopify App Merchant Finder | Vertical signal | Reverse-lookup tool: input an app name (Klaviyo, Recharge, Gorgias) and get the list of Shopify stores using it. |
| Doctor Directory Lead Finder | Vertical lead-gen | NPI Registry-grounded US doctor leads with Healthgrades/Vitals enrichment. Specialty + state filter across all 50 states. |
| GitHub Trending Scraper | Dev signal | Daily, weekly and monthly trending repos by language. Track emerging open-source projects and developer momentum as a startup signal. |
Additional actors covering Techstars, AngelList/Wellfound, 500 Global, Indie Hackers, Crunchbase News, NPM Package Stats, PyPI Package Stats, and Hugging Face model catalogs are in private testing and will flip public over the next few weeks. Watch the NexGenData store on Apify for new additions, or subscribe to the changelog.
Workflow example: building a VC-portfolio-overlap sales prospect list
A common ask from B2B SaaS sellers targeting late-seed to Series B startups: “give me every company backed by at least one of Founders Fund, Index Ventures or Lightspeed, deduplicated, enriched with a verified contact email, ranked by recency of investment.” Here is the end-to-end recipe using only tools in this category.
Step 1 β Pull the three portfolios in parallel. Trigger the Founders Fund, Index Ventures and Lightspeed portfolio scrapers via the Apify API or the web UI. Each returns a structured dataset with company name, website, sector, stage and partner attribution. Combined output is typically ~1,200 unique companies across the three firms. Total runtime: under five minutes if you fire all three in parallel.
Step 2 β Deduplicate on domain. Concatenate the three datasets, normalize the website field (strip protocol, www., trailing slash, lowercase), and dedupe. Co-investments β common at Series B+ β get collapsed but preserved as an array field listing every backing firm. This dedup pass typically removes 8β12% of rows and gives you a “co-invested by” signal that is genuinely useful for prioritization: a company backed by both Index and Lightspeed at the same round is a stronger signal than a single-firm bet.
Step 3 β Enrich with contact data. Feed the deduplicated domain list into the Company Enrichment Tool. For each domain you get employee count, industry, tech stack signals, LinkedIn URL and pattern-matched email addresses (founder@, ceo@, plus name-based guesses verified via SMTP probe). Drop any company with fewer than 10 employees if you only sell upmarket, or fewer than 50 for true enterprise targets. You can also drop any company whose tech stack already includes a competing product β instant disqualification rather than a wasted SDR cycle.
Step 4 β Add a recency signal. Sort by the “year_invested” or “stage” field from the original portfolio scrape. A company that just took its Series B six months ago is a far better prospect for an infrastructure SaaS sell than one that raised its Series A in 2018 and has since gone quiet. Pair with the Startup Funding Tracker to layer in any post-portfolio rounds that the VC’s own page has not yet been updated to reflect β Form D filings beat marketing pages to the news by days or weeks.
Step 5 β Push to your CRM. Pipe the final list to HubSpot or Salesforce via the Apify Hightouch integration or a simple webhook to your warehouse. Tag each record with the source VC firms so your AEs know whether to lead with the Founders Fund partner intro or the Lightspeed partner intro in their cold email. Add a custom property for “co-invested by count” so your sequencing logic can prioritize multi-firm bets.
Total cost for a 1,200-company sweep on Apify pay-per-result pricing: typically under $10. Same workflow on Crunchbase Pro plus ZoomInfo: $40K+/year in seat licenses. The data sources are the same β VC firms’ own public portfolio pages and public company registries β you’re just cutting out the aggregator middleman and owning the data pipeline end-to-end.
Variants of this workflow: swap in the Greylock + Bessemer + a16z portfolios for a different fund mix; add Sequoia and Peak XV for a US + India overlap list; filter Step 2 by sector to get only fintech, healthtech or developer-tool companies; or skip Step 3 entirely and pipe the raw deduplicated list into Apollo or Clay for their enrichment instead.
Related reading
- Map Any Vertical’s Competitive Landscape Using the YC Database (with code)
- How VC Associates Source Pre-Seed Deals Before Crunchbase Hears About Them
- Crunchbase Pro Alternative: 4 Free Sources for Real-Time Funding Data
- Crunchbase Killed Its Free API. Here’s How to Rebuild It (2026)
- Lead Enrichment Pipeline: From Domain to Full Company Profile (Free Stack)
- Track YC Demo Day Companies in Real Time (with code)
Related categories
- Market Intelligence Tools
- Regulatory & Compliance Tools
- Lead Generation Tools
- Financial Data Tools
- Public Registry & Institutional Tools
- Real Estate & Regional Tools
- Asian Business & Financial Tools
Frequently asked questions
How is this different from Crunchbase or PitchBook?
Crunchbase and PitchBook are aggregated databases sold under expensive seat licenses with strict scraping ToS. The NexGenData actors pull from the same public sources those databases also crawl (VC firm portfolio pages, accelerator alumni lists, SEC EDGAR filings, public news feeds) and return structured JSON you fully own. There is no seat limit, no API rate card, and the data lands directly in your data warehouse, CRM or spreadsheet at a fraction of the cost β typically pennies per thousand companies. The trade-off: you get raw structured data rather than the polished UI and curated firmographics layer the paid platforms add on top, but for most sourcing and prospecting workflows the raw data is exactly what you want.
How fresh is the data?
Each actor scrapes live on every run. VC portfolio pages typically refresh as soon as the firm publishes a new investment (often within hours of a round closing). Accelerator directories refresh as new batches go live. The Startup Funding Tracker polls SEC EDGAR daily for new Form D filings, which is often where you find rounds before the press release lands. Schedule any actor to run hourly, daily or weekly via the Apify Scheduler and you have a live data feed β no waiting on Crunchbase’s editorial queue.
Can I pipe results into HubSpot, Salesforce or my CRM?
Yes. Every actor returns clean JSON or CSV. The most common pattern is: run the actor on a schedule, push the dataset into your warehouse (Snowflake, BigQuery, Postgres) via webhook or Apify integration, then sync the deduplicated company list to HubSpot or Salesforce via reverse-ETL tools like Hightouch or Census. The Company Enrichment Tool is specifically designed to hydrate empty CRM fields (domain, employee count, industry, tech stack) before push, so your sales ops team gets fully-populated records rather than half-empty shells.
What about international VCs and non-US accelerators?
The current VC portfolio scraper lineup covers the US/EU corridor (Founders Fund, Lightspeed, Index Ventures, Greylock, Bessemer, a16z, Sequoia + Peak XV). The 500 Global directory covers MENA, LatAm, SEA, Korea, Taiwan, Vietnam and Eurasia regional cohorts. Index Ventures and Sequoia explicitly cover their European and Indian portfolios respectively. More European, Asian and emerging-market VCs are in the publishing pipeline β request additions by emailing the NexGenData team and they typically ship within two weeks.
Can I monitor a specific portfolio continuously?
Yes. Set up an Apify Schedule to run the relevant portfolio scraper daily or weekly. Each run writes a timestamped dataset, so a simple diff between yesterday’s and today’s dataset surfaces newly-added companies, status changes (e.g. ‘Active’ β ‘Acquired’) and partner reassignments. Pipe the diff into Slack or email for a near-real-time portfolio change feed. Many sourcing analysts run this exact setup against the a16z and Sequoia portfolios to get sub-24-hour notice of new investments.
What does this cost compared to paid databases?
Crunchbase Pro is around $1,000/user/year and PitchBook seats run $25K+/year. Most of these actors cost a fraction of a cent per company returned on Apify’s pay-per-result pricing. A full sweep of all five new VC portfolios (Founders Fund + Lightspeed + Index + Greylock + Bessemer β roughly 2,000 companies combined) typically costs under $5 and runs in a few minutes. You only pay for what you actually pull, with no minimum commitment, no seat fees, and no overage penalties. For most small-to-mid-size sourcing teams the entire annual data spend lands under $500.
Get started
Every actor on this page runs on Apify with pay-per-result pricing and a generous free tier. Pick the VC portfolio or startup directory closest to your sourcing thesis, run it once to validate the schema, then schedule it to refresh daily. If you need a custom build β a specific regional VC, an unlisted accelerator cohort, or a deeper enrichment pipeline that joins multiple data sources β get in touch and we’ll ship it for you, typically within two weeks of scoping.
Browse all NexGenData actors on Apify β
Recently added β hiring signals
- Workday Jobs Scraper β extract job postings from Workday-powered career sites.
Featured tool pages
Dedicated pages for our most-used lead-gen actors:
- Company Enrichment API β domain to firmographics
- Company Email Finder β verified emails by domain
- B2B Leads Finder β Apollo/ZoomInfo alternative
Factual public data — not financial or investment advice.