How SaaSafras data is made
SaaSafras is a research database on how software businesses find customers, price, and grow. This page says where every number comes from, what the labels on it mean, and what the data cannot support. If a claim on the site is not covered here, treat it as unverified.
v1-snapshot-2026-08-27. Next refresh: top 30 to 50 companies, dated per field.01What this is
One curated cohort, not a market. The database holds 140 companies (150 rows before deduplication) chosen for being interesting bootstrapped or mostly bootstrapped software businesses, assembled by one researcher. Funding is a per-row field (bootstrapped, partly-vc, vc, undisclosed): notion and calendly are VC-funded, carrd-co partly, and they stay in as the contrast cases. Every statistic on the site describes this list. None is a claim about SaaS in general.
Each company record has two layers. The prose profile (growth playbook, distribution moat, builder's takeaway, growth timeline) is written analysis. The structured fields (channels, category, growth stage, revenue, launch year, audience) are what cross-company analysis runs on. The prose came first. The structured fields were derived from it in August 2026 under written rules, and this page names those rules.
02Where the information comes from
The v1 dataset was collected in February 2026 from public material: founder interviews and podcasts, company blogs and about pages, pricing pages, acquisition announcements, open-metrics pages, and founder posts on X, Indie Hackers, and Hacker News. Candidates were found by hand and through Product Hunt and Indie Hackers listings. Extraction and first-pass classification used language models with human review; the analysis prose was written by the researcher.
Every record carries a Sources field. Its current state, measured on 2026-08-27:
| Sources field | Rows | Meaning |
|---|---|---|
| Filled | 150 of 150 | Every row names its sources. |
| Contains a link | 27 of 150 | The rest name sources in prose ("founder interviews", "company blog") without a URL. |
| Link points at the revenue figure itself | 0 of 150 in v1; 15 of 30 after the refresh | No v1 revenue number is backed by a link to the statement it came from. The top-30 refresh (2026-08-27) added dated, linked revenue observations; 15 of those 30 companies now have at least one (section 06). |
| Starter Story claims layer | 166 companies, 184 founder-stated figures; 5 confirmed | Revenue figures founders state on camera in Starter Story interviews (437 transcripts, 2026-08-27). Every row is labeled Founder reported with the video link, the quote, and the date. These are claims, not verification: a second source was sought for all 166; 5 were confirmed, 83 have a real product and founder but no independent dollar figure, 5 were contradicted, 73 could not be checked. The four counts are published with the rows; a claim never becomes a verified figure by being repeated. |
That is the honest starting point. The refresh described in section 11 replaces prose sources with dated links, field by field, starting with revenue.
Growth timeline steps already carry a per-step tag: documented when a source states the event, inferred when the researcher deduced it. Across all 150 rows: 585 documented steps, 8 inferred, 1 documented with a stated caveat.
03How companies are selected
By interest, not by sampling. A company entered v1 because it was a bootstrapped or mostly bootstrapped software business with enough public material to write a growth playbook from. Launch years run 1996 to 2025. There is no random sample, no size threshold, and no attempt to represent a category in proportion to its size.
Three consequences follow, and they are stated again in section 12: no dead companies, no denominator, and the researcher's taste is in the selection.
Deduplication
Separate collection passes picked up 8 companies more than once, as 17 rows under different slugs: Ghost, Carrd, Transistor.fm, Lemon Squeezy, Plausible, Nomad List, Remote OK, and Fathom. The rule: group rows by company name, keep the row carrying a numeric MRR; if more than one does, keep the row with more channels, then the higher MRR. The mapping lives in a versioned file and every analysis reads it. The source CSV is never edited to remove a duplicate, so provenance survives.
One company, one id, kept across renames and pivots. Old slugs resolve as aliases.
04Provenance labels
Four labels, applied per field, not per company. A company can have a founder-reported price and an estimated MRR in the same record.
| Label | Means | Trust |
|---|---|---|
| Founder reported | The founder stated it directly: interview, podcast, public post, DM, or a form submission to SaaSafras. | Highest for subjective facts (motivation, story). Not proof for numbers; founders round, and they report peaks. |
| Publicly reported | Stated in a public, citable source not authored by SaaSafras: press, filing, acquisition announcement, the company's own site, socials, or open-metrics page. | Default for anything pulled from a URL that is not a founder conversation. |
| Estimated | SaaSafras inferred the value from indirect signals: pricing-page math, traffic tools, comparable companies, a stated range. | Always carries a note on how. Order of magnitude, not a figure. |
| SaaSafras classification | SaaSafras assigned a value from its own vocabulary: category, channel, growth stage, business model, audience. | The rule is public (sections 8 and 9). Disagree with the rule, not the row. |
Status 2026-08-27: classification fields (channels, category, growth stage, audience) are labeled by construction. Revenue, pricing, and status fields carry a label only after the refresh reaches them; until then they are unlabeled v1 values and should be read as Estimated.
05How revenue figures are handled
The revenue column is Normalized MRR (USD): a monthly figure in US dollars. 111 of the 140 companies carry a numeric value; 29 have none and are excluded from every revenue table rather than filled in. All figures are USD; the cohort contains no non-USD reporting.
v1 mixed metric types in that one column. The same cell can hold current MRR, an annual figure divided by twelve, lifetime revenue divided by months in business, or an acquisition price. Example: Laravel Forge shows $333,333 per month, which is $40M lifetime over 120 months, not an observed MRR. This is the single largest quality problem in the dataset and the first thing the refresh fixes.
Rules going forward
- Every revenue figure gets a
revenue_metric_type: current MRR, stated monthly revenue (a per-month figure not called MRR), ARR, stated annual revenue (not called ARR), launch-window revenue (a total inside 30 days of launch), lifetime revenue, acquisition price, or run-rate estimate. Figures of different types are never compared or averaged. Funding (bootstrapped, partly VC, VC, undisclosed) is its own field, kept out of the status column. - Every figure gets a date (
as_of) and a provenance label. A figure with no date is a v1 figure and reads as "at some point before February 2026". - USD is assumed where the source writes "$" and names no currency. Where the source names one, the record says so; where the figure is spoken in a non-US interview with no currency stated, the record carries a currency-unstated note.
- Revenue history is append-only. A new figure adds an observation; it never overwrites the old one. Old figures stay visible with their dates.
- Until a company's figures have been through the refresh, treat every dollar amount on the site as an order of magnitude.
06What verified means
Verified means all of the following are true for one field on one date: a link to a source that states the value, the metric type recorded, the provenance label recorded, and a human who read the source on a named date (last_verified_date).
A field that is missing any one of those is not verified, whatever the prose around it says.
Status 2026-08-27: the test has been applied to 30 companies (the top-30 refresh, 229 dated observations, every cited revenue source opened by a human). 15 of the 30 have at least one revenue field that passes: Profitwell, Doist, Ghost, Nomad List, Formula Bot, Carrd, TinyPilot, Statuspage, Baremetrics, Plausible, Tally, Photo AI, Balsamiq, TypingMind, Bannerbear. The other 15 have no source that states a current figure, or the only sources repeat each other with no founder or filing behind them; their v1 figures are withdrawn, not replaced. No company outside the 30 has been tested, so the v1 figure on any other row is unverified by this definition.
Verified is not "true". It is "we can show you where it came from and when we looked". A founder can be wrong on a podcast; the label tells you the chain, not the truth.
07What an estimate is
An estimate is a value SaaSafras produced when no source states it outright. Every estimate names its method in the record. The methods in use:
- Pricing-page math: stated customer count multiplied by the entry price, or a stated ARR divided by twelve.
- Range midpoint: a founder says "$80K to $110K"; the record stores 95,000 with the range noted, never the midpoint alone.
- Comparable: a company of similar size and model with a reported figure. Weakest method; used only when the alternative is a blank and the record says so.
- Historical carry-forward: a figure reported in an earlier year with no update since. The date on the figure is the reported date, not today.
Estimates are never presented with more precision than the method supports. A comparable-based MRR reads as a band, not a number.
08How channels are classified
Acquisition channels are a 12-value, multi-valued field: seo, paid, outbound, community, social, product-led, partnerships, marketplace, audience-led, content, word-of-mouth, launch-platform. Multi-valued because the data forced it: the v1 prose named 593 channel mentions across 150 companies, 3.45 per company, and a single value would have discarded three quarters of the column.
Each value has a written definition and eight disambiguation rules (community is a place with other people's threads; social is broadcasting to your own followers; content is the artifact pulling, SEO is the ranking pulling; and so on). The 307 distinct phrases in v1's channel prose were mapped against that taxonomy, cross-checked by an independent keyword mapper, and 96.5 percent of mentions resolved. Unmappable phrases were recorded as unmapped, not guessed: one company (a customer count where a channel should be) has zero channels on purpose.
The taxonomy grows only when measured data shows a value with no honest home, not when a category sounds plausible. Three of the twelve (content, word-of-mouth, launch-platform) were added that way, each a top-5 phrase by frequency with no place in the original nine.
Channels are recorded after success, from what founders said about it. That is a description of what winners report, with the hindsight that carries. See section 12.
09Categories and growth stage
Category
Nine buckets plus Other: Marketing & Growth Tools, Analytics & Monitoring, AI & Creative Content Tools, Developer Tools & Infrastructure, E-commerce & Shopify Tools, No-Code Builders & Content Publishing, Productivity & Collaboration, Vertical/Niche Business Software, Consumer Apps & Games. The 128 distinct v1 category strings map to these through a published table. AI & Creative holds generative products only; AI-powered support, productivity, and infrastructure tools sit in the bucket their buyer would look in. Other holds one company (hardware). Every assignment is a SaaSafras classification.
Growth stage
Four values: build, launch, operate, scale. v1 described each company's stage in a paragraph; the leading word of that paragraph is the stage. Exited and acquired companies fold into scale, and the exit itself is kept in a separate exit_note (acquired, sold, at exit) so the signal is not lost. Current distribution across 140: build 8, launch 15, operate 50, scale 67.
Audience
b2b, b2c, mixed, or unclear, by a keyword rule over the audience prose. Unclear is a real value; it is not resolved by guessing.
The AIBuilder flag
25 v1 rows carry an AIBuilder flag whose original definition was not recorded. 23 of the 25 launched 2021 to 2025, and the flag behaves like a recency marker. It is not the AI-native cohort. That cohort is being built separately with its own criteria, and no before-and-after comparison is published until it exists.
10Conflicting sources
When two sources disagree, both values are kept, each with its date, metric type, and label. SaaSafras does not average them and does not pick one silently. The record shows the conflict; the analysis uses the figure whose metric type matches the question being asked.
Four genuine conflicts were known in v1; each was resolved on 2026-09-01 as shown below:
| Company | Conflict |
|---|---|
| Zenvoice | $236 per month (Oct 2025) and "$1.3K per month peak" (Mar 2024 launch month) are the same metric type at different dates, not a genuine conflict; resolved 2026-09-01 by re-dating both, and TrustMRR's Stripe-verified figure puts current MRR at $0 as of Sep 2026. |
| 17hats | MRR $75,000 in the numeric column had no source at any era; withdrawn 2026-09-01. The only founder-attributed figure on record is a 2015 TechCrunch report of $2M first-year revenue, a different metric. |
| TinyPilot | 95,000 was stored as a point inside a stated $80K to $110K range, without the range marked; resolved 2026-09-01 by marking the range in flags, value unchanged as the midpoint. |
| Laravel Forge | $333,333 per month was $40M lifetime over 120 months, not an observed MRR; withdrawn 2026-09-01, the lifetime figure retained in revenue_fragment_raw. |
A formal precedence rule (which label wins when types match) is deliberately not published yet, because until every figure has a metric type there is nothing to arbitrate. It ships with the refresh.
11When records were last updated
| Event | Date | What changed |
|---|---|---|
| v1 collection | 2026-02 | 150 rows, 27 columns, prose profiles, growth timelines with per-step tags. |
| v1 snapshot frozen | 2026-08-27 | Tag v1-snapshot-2026-08-27. The archival copy; never edited. |
| Deduplication + normalization | 2026-08-27 | 141 canonical companies; channels, category, growth stage, audience derived under the rules above. No revenue figure was changed. |
| Top 30 refresh | 2026-08-27 | 30 companies, 229 dated observations across revenue, pricing, website, status, and major GTM developments, each field dated and labeled, every cited revenue source opened. 15 of 30 companies pass the verified test (section 06). |
| Starter Story claims layer | 2026-08-27 | 437 interview transcripts parsed; 184 founder-stated revenue figures across 166 companies landed as dated, labeled claims (section 02), plus 170 second-source observations (status, website, founder, pricing, revenue) from the verification pass. 5 of 166 confirmed, 5 contradicted; the rest are labeled as unverified claims. No v1 figure was changed. |
| Corrections pass | 2026-08-27 | 20 v1 figures superseded or withdrawn as dated observations (Pingdom $103M to $67M per filing; Statuspage $30M+ to $18.3M + $3.3M per filing; Nomad List portfolio figure withdrawn; 12 unsourced MRR/ARR figures withdrawn). v1 values were not edited; each correction quotes the v1 value beside it. |
| Website + conflict resolution pass | 2026-09-01 | 97 of 111 missing website_url cells filled (curl-verified web pass plus disk-join from existing observations), 14 recorded as punts (domain squatted, ambiguous, or unreachable). Four genuine revenue conflicts resolved: 17hats and Laravel Forge unsourced or derived MRR withdrawn (mrr_usd cleared, flagged), TinyPilot range marked, Zenvoice re-dated as same-metric different-month figures. These are the only v1 mrr_usd values edited to date. |
| Tweet Hunter dedupe | 2026-09-02 | tweet-hunter and tweethunter confirmed one product (tweethunter.io, caught by identical website in the 09-01 web pass) and merged; 141 canonical companies became 140. The kept row carries the revenue figure; the dropped row's channels were folded in. |
| Companies 31 to 50 | pending | Same refresh, same test. |
A record with no per-field date has not been refreshed since v1 collection. Once a company has been refreshed, its profile shows the date beside each figure, and the old figure stays in its history.
12What this data cannot say
- Nothing about the market. 140 companies chosen for being interesting successes. "Paid ads at 1 percent" describes the list, not SaaS.
- Nothing causal. Channels were recorded after the outcome, from what founders said. Associations here are descriptions of what winners report.
- Nothing about failure. No dead companies. There is no denominator.
- Nothing precise about revenue. Metric types are mixed and unlabeled in v1; 29 companies have no figure. Order of magnitude only, until the refresh.
- Nothing about the AI-native era yet. The flag in v1 is a recency marker with an unrecorded definition. The comparison this database exists to make waits for a cohort collected under stated criteria.
Every published table reproduces from scripts in the dataset repository against the frozen snapshot. Changing a classification means editing the published mapping and re-running, never editing a row by hand.
13Corrections
If a figure about your company is wrong, say so. A founder correction becomes a Founder reported observation with today's date; the old value stays in the record's history with its own date, so readers can see both. Corrections that come with a link get the Publicly reported label instead.
Send corrections to [email protected].