Validation Report

Your idea: "An app that helps students find part-time jobs faster"

Riskiest assumption

You can reach these users through a repeatable, low-cost channel — not just your first-degree network.

If it's wrong: CAC balloons, growth stalls after friends-and-family, and you burn months on ads that don't work.

Free alternatives check

Can the target user already get most of this outcome today for free?

Partial free alternativeDoing it manually once a monthA free ChatGPT promptAn existing free community

Free options cover part of the outcome. You need evidence that the missing part is the part users actually care about — shown by behavior, not opinions.

Do this first

Concierge test: deliver the outcome manually for 20 freelance designers

7 days · $0–50 · one test, one decision.

Assumption it tests

You can reach these users through a repeatable, low-cost channel — not just your first-degree network.

Why this test

It is the cheapest way to observe real behavior instead of collecting language. You deliver the outcome by hand, so nothing gets built until someone acts.

What counts as real evidence

Real evidence is an action, not a sentence: a user completes the core action twice without a reminder, shares their real data or workflow, schedules a follow-up, pays, or brings in the budget owner. Compliments do not count.

Keep if

At least two independent signals appear, with at least one behavioral or economic — for example 5 of 20 users repeat the core action unprompted AND at least one pays, preorders, or involves a budget owner.

Pivot if

Users engage but only at language or commitment level — they agree, they reply, they book calls, but nobody changes a workflow or spends anything. The pain is real; this framing or segment is not the one.

Drop if

You reached 20 of the right users through the right channel and nobody moved past “cool idea” — and the free alternative kept winning.

Secondary options (run only after the first test)
10 problem interviews

15-minute unscripted calls. Ask about the last time the problem happened. Never pitch.

Why not first: Great for context, but interviews mostly produce language-level signal.

Pre-sale / deposit test

Offer a founding-user plan with money upfront, fully refundable.

Why not first: Strongest signal, but usually needs the concierge test first to know what you are selling.

Fake-door landing page

One-page site with the promise and a single CTA. Drive 200 targeted clicks from one channel.

Why not first: Measures interest, not behavior. Easy to pass, weak as evidence.

Frozen kill threshold

Within 7 days, 5 of 20 target users must complete the core action twice without a reminder. If not, pivot or change the segment.

Before the test starts, Needly helps you define the kill threshold. Once the test begins, the rule is frozen. This prevents founders from moving the goalposts after weak results.

Signal strength
LanguageThey say the pain is real.
CommitmentThey schedule a follow-up, share data, or give access.
BehaviorThey change a workflow or repeat the core action unprompted.
EconomicThey pay, preorder, sign an LOI, or bring in the budget owner.

You have not run the test yet, so the only signal available is language. Language alone never validates an idea.

At least two independent signals, one of them behavioral or economic.

Decision confidence
Medium

Confidence stays low until real behavior is observed. This report ranks what to test first — it does not predict whether the idea will succeed.

Recommendation
Keep going

Sell the concierge version to 3 people at full price before writing code.

Four receipts behind this recommendation

Assumption being tested

You can reach these users through a repeatable, low-cost channel — not just your first-degree network.

Precommitted threshold

Within 7 days, 5 of 20 target users must complete the core action twice without a reminder. If not, pivot or change the segment.

Raw evidence collected

No evidence collected yet. Run the recommended test, then log what actually happened — not how it felt.

What would reverse this decision

Zero users move past language-level signal, or the same outcome is already available free.

Idea failed or test failed?

Sometimes the idea is weak. Sometimes the test was weak. Needly separates both.

The idea is weak if
  • You reached the right users and they still did nothing.
  • They preferred an existing free alternative (Free habit trackers).
  • No one moved past saying the pain is real.
The test was weak if
  • Fewer than 15 target users actually saw the ask.
  • The channel reached a different segment than the one you defined.
  • The ask was vague, so no one had a concrete action to take.

Don't kill an idea on a broken test. Rerun the test properly before you decide.

Go deeper

Everything behind the score. Open only what you need.

Evidence behind the score
Needly Score and the five weighted drivers
Needly Score
78/ 100

Strong signal. Score 78/100 — the pain is real. Test before building.

The score ranks what to test first. It does not predict success.

Pain intensity
Weight: 30%
90/100

Pain shows up as "critical" — based on frequency, urgency, and cost of the workaround described in similar problem spaces.

Customer clarity
Weight: 20%
50/100

Target user is specific: "Freelance designers". You can name them, find them online, and reach them without paying for ads.

Testability
Weight: 15%
83/100

A meaningful test can ship in a weekend using a landing page + 10 conversations — no code required.

Competitive wedge
Weight: 20%
68/100

3 known alternatives (Free habit trackers, …). Wedge exists but is narrow — differentiation must be sharp.

Willingness to pay
Weight: 15%
79/100

Users in this segment already pay for adjacent tools, so willingness-to-pay is plausible.

Demand signals
Weak signals vs strong signals to look for

A waitlist can be misleading. Interviews can be biased. Compliments are not validation. Needly looks for stronger signals: repeated pain, urgency, willingness to pay, follow-up behavior, and clear access to the target customer.

Weak signals
  • Compliments
  • “Cool idea”
  • Generic positive feedback
  • Waitlist signups with no action
  • Survey answers with no commitment
  • Users agree the problem exists but keep their current workaround
Strong signals
  • Users schedule a follow-up
  • Users share real data or workflow details
  • Users change their current workflow
  • Users repeat the action without reminders
  • Users pay, preorder, leave a deposit, sign an LOI, or involve the budget owner
  • Users introduce another person with the same problem
Pain level
CriticalUsers are actively losing time or money every week and complain publicly.
Competition
Existing alternatives and where the wedge is

An app that helps students find part-time jobs faster — currently, the people you're targeting solve this with fragmented, manual tooling that doesn't scale past a few uses.

  • Free habit trackers
  • Journaling apps
  • Just willpower
Target users
  • Freelance designers
  • Independent consultants
  • Small agency owners
Monetization potential
Willingness to pay and what would prove it
79/ 100

Users in this segment already pay for adjacent tools, so willingness-to-pay is plausible.

Payment intent is the only signal that survives contact with reality. Ask for a preorder, a deposit, or a signed letter of intent — not a “yes, I'd use that”.

Testability
The fastest experiments you can run this week
10 problem interviews in 5 days

Book 15-min calls with target users. Ask about the last time the problem happened, never pitch.

Green light: ≥ 7 out of 10 describe the problem unprompted in the last 30 days.
Fake-door landing page

Ship a one-page site with the headline and a 'Get early access' email capture. Drive 200 clicks from 1 channel.

Green light: ≥ 5% email conversion from cold traffic.
Concierge MVP for 3 users

Deliver the outcome manually — no product, just you in a shared doc — for 3 real users.

Green light: At least 1 asks 'how much?' before you offer.
Landing page test

Built for people who care about part-time.

Ship a one-page site with the headline below and an email capture. Drive 200 clicks from one channel (Reddit, X, or a niche newsletter). Under 3% opt-in = the promise isn't landing.

Key assumptions
What must be true — and the main risks
  1. 01You can reach the target users cheaply (< $50 CAC).
  2. 02The current workarounds are painful enough to switch.
  3. 03There is a repeatable acquisition channel you actually enjoy.
Main risks
  • Users describe the problem but keep tolerating cheap workarounds.
  • Crowded space — Free habit trackers already covers 80% of the job. You need a sharp wedge.
  • Retention past week 2 is unproven — novelty may inflate early signal.
Why the weakest assumption matters

Without distribution, even a great product dies. Most ideas fail on channel, not on product.

How you'll know

One channel delivers 50 targeted visitors for < $30 in a 48-hour test.

Interview insights
5 non-leading questions to run this week
  1. 1.Walk me through the last time you ran into this. What did you do?
  2. 2.What did you try before? Why didn't it stick?
  3. 3.How much time or money is this costing you per month?
  4. 4.If we removed this problem tomorrow, what would change in your week?
  5. 5.Who else on your team feels this pain? Who cares the most?

Never pitch during an interview. Ask about the last time the problem happened, then listen for what they already tried.

Day-by-day plan for the recommended test
7 days, timeboxed, one decision at the end
Hypothesis

Freelance designers feel this problem often enough — and painfully enough — to sign up for a solution within 48 hours of seeing it.

  1. 01
    Day 1 — MonSharpen the promise
    60 min

    Write a one-sentence landing headline and the single CTA. Rewrite it 5 times until a stranger understands it in < 10 seconds.

  2. 02
    Day 2 — TueShip the fake door
    90 min

    Build a one-page site (Framer, Carrd, or a simple TSX route) with the headline + email capture. No product, no roadmap.

  3. 03
    Day 3 — WedLine up 10 interviews
    60 min

    DM 30 freelance designers on the platform they actually hang out on. Ask for 15 minutes. Aim to book 10.

  4. 04
    Day 4 — ThuDrive 200 targeted clicks
    90 min

    Push the landing page in 1 channel (subreddit, niche newsletter, Slack). Track opt-in rate, not vanity metrics.

  5. 05
    Day 5 — FriRun 5 interviews
    90 min

    Ask about the last time the problem happened. Never pitch. Record and take verbatim quotes.

  6. 06
    Day 6 — SatRun 5 more + tally signal
    90 min

    Finish interviews. Count opt-ins, unprompted mentions of the pain, and 'how much?' asks.

  7. 07
    Day 7 — SunDecide: keep, pivot, or drop
    45 min

    Compare the numbers below against the success and kill signals. Write a 1-page decision memo. No 'maybe'.

Decision criteria
The rules Needly applies when reading your evidence

Before you test, Needly helps you define what evidence is enough to continue — and what evidence means you should stop or pivot.

  • Fewer than 3 out of 15 target users agree to a follow-up.Pause or pivot
  • Nobody shows payment intent — no preorder, no deposit, no budget owner.Pause or pivot
  • Users say “sounds interesting” but take no action afterwards.Weak signal
  • Multiple users describe the same urgent pain, unprompted.Continue
  • People are willing to preorder, pay, or introduce you to others.Double down

A waitlist can be misleading. Interviews can be biased. Compliments are not validation. Needly looks for stronger signals: repeated pain, urgency, willingness to pay, follow-up behavior, and clear access to the target customer.