Is My Startup Idea Good? Use This Evidence-First Scorecard
Use an evidence-first startup scorecard to examine demand, urgency, competition, differentiation, timing, business model, GTM, founder fit, and risk.
A good startup idea gives a specific customer a strong reason to change behavior, and it gives the founder a credible way to deliver and distribute that change. You cannot discover that from novelty alone. The useful question is not “Does this sound clever?” but “Which parts of this idea are supported by evidence, and which part could still kill it?”
This evidence-first scorecard examines nine dimensions: demand, urgency, competition, differentiation, timing, business model, go-to-market, founder fit, and execution risk. It deliberately does not produce a magic total. A precise number built on guesses is less honest than a clear map of what you do not know.
A good idea is not the same as a clever idea
An idea can be technically impressive and commercially weak. It can solve a real problem for people who cannot buy. It can enter a growing market without a reason to win. It can attract signups while failing to create repeated value. Conversely, a plain solution to an expensive, recurring problem can be an excellent starting point.
The scorecard therefore asks for evidence beside every judgment. Use four evidence states:
- Assumption: the claim is plausible but unsupported.
- Observation: target customers consistently describe or demonstrate it.
- Commitment: target customers spend time, reputation, access, or budget.
- Behavior: the target action actually happens, ideally repeatedly.
The states are not interchangeable. A market report can support category context but cannot show that your chosen buyer will switch. Ten enthusiastic interviews can reveal a workflow but cannot establish retention. Revenue from one custom consulting project can establish payment but not repeatable product demand.
YC's current interview guide asks founders about users, acquisition, growth, retention, unit economics, buying resistance, and surprising behavior. Those are useful questions because they force the idea to meet an operating reality. The scorecard below applies the same discipline before or just after launch.
The nine dimensions
1. Demand
Diagnostic question: Does a clearly defined group already experience the problem and take action around it?
Evidence requirement: Recent examples from the target segment: attempted workarounds, money spent, people assigned, tools combined, deadlines missed, or risks accepted. Search activity and community discussion can show context, but direct workflow evidence is stronger.
Warning sign: The problem appears only after you explain the solution. People agree it would be nice, yet cannot describe the last time it mattered.
Demand is not the size of a broad category. “The creator economy is large” says nothing about whether independent course authors need your refund-prediction tool. Narrow the customer and situation until the behavior can be observed.
2. Urgency
Diagnostic question: What makes the customer act now instead of tolerating the problem for another year?
Evidence requirement: A trigger with consequences: a compliance deadline, lost revenue, repeated manual work, a new role, an incident, a budget cycle, or a change in an upstream platform. Ask what happened the last time the trigger appeared and what the customer did next.
Warning sign: The value is expressed as general improvement with no clock, owner, or consequence. “Save time” is weak until you know whose time, how often, and what that time displaces.
Urgency can vary within the same customer type. A bookkeeping problem may be mild for a solo consultant and acute for an agency closing ten client books at month-end. Segment around the moment of pain, not demographics alone.
3. Competition
Diagnostic question: What does the customer use today, including doing nothing, and why is that alternative acceptable?
Evidence requirement: A current alternative map built from product pages, pricing, documentation, reviews, and customer workflow. Record the job each alternative performs, its switching cost, and the reason customers keep it.
Warning sign: “We have no competitors.” That usually means the search used your category label rather than the customer's problem. A spreadsheet, an assistant, an agency, an adjacent feature, or tolerated pain can all compete.
Competition can support demand because it shows people already allocate effort or money. It also raises the standard for differentiation. Do not confuse a crowded list of products with proof that your chosen wedge is reachable.
4. Differentiation
Diagnostic question: Why would the target customer switch from the current alternative, and can they notice that difference before paying a large switching cost?
Evidence requirement: Comparative tests, lost-deal notes, observed gaps, design-partner commitments, or usage showing that a narrow promise matters. The difference should connect to an outcome: lower risk, faster completion, unique access, easier adoption, or a better fit for a neglected workflow.
Warning sign: The difference is a feature adjective: smarter, simpler, AI-powered, all-in-one, or cheaper without an economic reason the advantage can persist.
Differentiation is not permanent. It is an initial reason to choose you. A credible distribution advantage or deep access to a customer workflow may be more defensible than another generated feature.
5. Timing
Diagnostic question: What changed recently that makes adoption possible or necessary now?
Evidence requirement: A verifiable shift in regulation, technology cost, platform capability, customer behavior, workflow, or distribution. Then show how the shift affects the target customer rather than citing a trend headline.
Warning sign: Timing is justified only with “AI is growing” or another broad narrative. A trend can create attention while making a category noisier and easier to copy.
Good timing also includes readiness. A technically feasible product can be too early if data is unavailable, trust is absent, or the buyer has no process for adopting it.
6. Business model
Diagnostic question: Who receives the value, who authorizes payment, and can the price support delivery and acquisition?
Evidence requirement: Buyer conversations, existing budget behavior, preorders, paid pilots, or early unit economics. Separate the user from the buyer and the budget owner from an enthusiastic champion.
Warning sign: Pricing begins with competitor averages or preferred revenue rather than the economic value and purchasing process. Another warning is a high-touch service hidden behind low self-serve pricing.
Strategyzer's Testing Business Ideas separates desirability, feasibility, and viability. Payment is valuable evidence, but one payment does not prove that margins, support, and acquisition work at scale.
7. Go-to-market
Diagnostic question: Can you name and reach the first 20 plausible customers through a channel that fits how they discover and buy?
Evidence requirement: Actual response, referral, conversion, or sales-cycle data from a specific channel. For a pre-launch idea, demonstrate access: community credibility, partner relationships, an audience, outbound lists with replies, or a workflow where the product can be encountered.
Warning sign: The plan is “content, SEO, and Product Hunt” without a reason the target customer is there, a message, or a conversion step.
A good market with no reachable wedge is still a bad starting position. Early go-to-market can be manual. The question is whether you can create learning and transactions, not whether the channel is already infinitely scalable.
8. Founder fit
Diagnostic question: Why are you unusually able or willing to understand, reach, and serve this customer for years?
Evidence requirement: Lived workflow experience, customer access, domain knowledge, technical leverage, sales ability, trusted relationships, or a demonstrated learning speed. Passion is relevant but not sufficient.
Warning sign: The founder chose the idea only because it is technically interesting or because a market report looks large. Another warning is contempt for the customer, sales process, or unglamorous support work.
Founder fit is not destiny. Gaps can be learned or complemented. The scorecard should expose the gap early enough to decide whether acquiring that capability is realistic.
9. Execution risk
Diagnostic question: Which technical, legal, operational, trust, or dependency risk can prevent the promised outcome even if customers want it?
Evidence requirement: A proof of concept for the critical technical path, expert review where regulation or safety matters, dependency checks, cost tests, and a manual walkthrough of the end-to-end service.
Warning sign: The demo works only on curated examples, delivery depends on an unstable third-party platform, or the product handles sensitive decisions without a credible reliability and escalation model.
Execution risk is not a reason to prefer easy products. It is a reason to test the hardest dependency before polishing the surrounding experience.
How to read a weak score
Do not average the nine dimensions. Find the weakest dimension that is both essential and poorly evidenced. One fatal constraint can dominate eight strengths.
Classify each dimension:
- Supported: evidence is strong enough for the next investment.
- Promising: observations are consistent, but a behavioral test is missing.
- Exposed: evidence contradicts the current idea or a critical dependency remains untested.
- Unknown: you have not looked outside your own reasoning.
Then inspect evidence quality. A “supported” demand judgment backed by friends' opinions should be downgraded. A “promising” founder-fit judgment backed by years in the workflow may deserve more confidence. This is why the ledger matters more than the label.
Weak does not automatically mean stop. It tells you what experiment deserves priority. If urgency is weak, recruit around the trigger and ask about recent consequences. If differentiation is weak, show alternatives side by side and ask for a costly next step. If execution risk is weak, build a technical spike, not a homepage.
What this self-check cannot tell you
This article cannot give you a real market verdict. It cannot run live competitor research, interview your customer, observe payment, or measure retention. It also does not reproduce RoastIdea's full startup scorecard. It is a structured self-check designed to catch unsupported confidence.
The dimensions are decision lenses, not a validated statistical model. No total can produce a defensible “72% chance of success.” Startup outcomes depend on changing markets, execution, learning, and events the scorecard does not observe.
Be especially skeptical when an AI fills every field fluently. Language models can propose competitors, risks, and experiments, but fluency is not evidence. A claim should link to a source or remain labeled as an assumption. For the full boundary, see what AI can and cannot prove about a startup idea.
Choose the next unknown to research
Convert the weakest consequential dimension into one falsifiable question. Then choose the smallest test that reaches the required evidence:
- For demand, investigate recent behavior and existing workarounds.
- For urgency, recruit around a real trigger.
- For differentiation, ask target users to choose or switch.
- For business model, request a budgeted commitment.
- For go-to-market, try to acquire conversations through the proposed channel.
- For execution risk, test the hardest dependency end to end.
Predeclare what result means continue, revise, or stop. Our guide to seven business-idea experiments helps match the test to the evidence you need.
The scorecard has done its job when it makes one unknown impossible to ignore. A good idea is not the one with the most attractive boxes. It is the one whose largest risks can be named, tested, and survived before they consume months of building.



