Guide · updated 2026-09-09
How to choose a QA outsourcing vendor
Selection starts before any call. Write down the two requirements that can disqualify a candidate, then name the public artefact that settles each one. The check that removes the most names is commercial rather than technical: ten of the twelve companies here publish a project threshold on a directory profile, and that one field decides fit without a meeting. Capability questions are worth asking only of the candidates that clear it.
What a scorecard has to do
A scorecard turns a preference into a question with an answer that can be filed, so two candidates are compared on the same evidence and the gap becomes a number. The eight dimensions behind the ranking work as a starting frame; the weights and the 0-100 scales sit on the methodology page, and every figure quoted below comes from the company files listed on the sources page.
Copy the frame, then change the weights to match your constraint. A team shipping a medical device weights governance and regulatory evidence first. A team with an eleven-week release window weights start time and exit terms instead. The weights matter less than the rule that each row closes on an artefact rather than on a conversation.
The scorecard: ten rows and what closes each one
| Criterion | What to ask for | What a solid answer looks like | What a weak answer looks like |
|---|---|---|---|
| Automation you can inspect | Two cases with a before-and-after figure, and the CI service the suite runs in | Coverage share or regression run time for a named product, with the pipeline named | A tool list, plus a claim that automation is part of every project |
| A certificate a stranger can check | Certificate number, registrar, expiry date, and the scope wording | All four in writing, or a register entry you can open without help | A logo, a screenshot, or a standard named without its registrar |
| Data handling in writing | The DPA, the sub-processor list, and whether production data is ever copied | A named controller entity, a stated transfer mechanism, masked or synthetic test data by default | GDPR compliance asserted on the privacy page, no document attached |
| Regulated-industry evidence | One finished engagement under the rule that applies to you | The regulation named, the product area named, the handover documents listed | An industry tag list with no case behind any entry |
| Start time to first engineer | Days from signature to first engineer working, and what has to be ready on your side | A day count committed in writing, with the access list attached | "As soon as you are ready", with no dependency list |
| Exit terms | Notice period, ownership of test code, and the account the repository sits under | Work product assigned to the client, a repository under your organisation, notice in days | Assets held in the supplier's tooling, exportable on request |
| Spread of the review record | Which directories carry a count, and the dates of the three newest entries | Counts on two or more platforms, entries inside the last twelve months | One platform holds the whole record, or a profile carries no count |
| Named buyers behind the cases | A case where the client is named, and a call with that client | A named buyer, one delivery figure, and a call in the diary | Client logos beside anonymous case texts |
| Price structure | Rate by role, what it includes, what is billed on top, the monthly floor | A rate card by seniority, with pass-through costs itemised separately | One blended figure, everything else left to the invoice |
| People and continuity | Engineers available on your stack, overlap hours, replacement terms | Named engineers, overlap in your working hours, notice and handover days | Total headcount, plus an assurance that a bench exists |
Score every candidate on the same rows and record the artefact beside each score; a row closed by a sentence in a proposal is not closed. Rate structures are covered in the QA outsourcing pricing guide.
Eight signals you can read before the first call
Half the scorecard can be filled in from public pages, and the faster reading is of what is missing rather than what is claimed. Each signal below appears in the material collected for this ranking, checked 2026-09-06.
| Signal | Where you see it | What it tells you |
|---|---|---|
| A certificate named without a number or a registrar | Service or About page on the supplier's site | ImpactQA names ISO 9001 and ISO 27001 on a managed-testing page with no certificate number or issuing body, so the audited scope cannot be established before contracting |
| A register entry held by the body that owns the standard | The standard owner's own directory | The TMMi Foundation register carries Kualitatem's Level 5 appraisal with a certificate identifier, an expiry in 2027 and Planit Testing named as the appraising provider. Among the twelve, it is the only entry that opens without a request |
| No certificate published at all | Certifications or About page, where the section is absent | Neither Cleverix nor QASource publishes a certificate, so every assurance question moves into the contract |
| A review record confined to one directory | Clutch, G2 and GoodFirms profiles side by side | iBeta Quality Assurance carries three reviews on one directory, with no G2 or GoodFirms count in public sources |
| Case studies that keep the buyer anonymous | The supplier's case library | Nine of the ten cases QualityLogic publishes name no client, so a reference call has to do the work the missing name would have done |
| A project threshold printed on the profile | The "Min project size" field on a directory profile | QASource publishes a $25,000 project floor, which settles fit for a smaller budget without anyone taking a call |
| No rate anywhere on the supplier's own pages | Pricing or services navigation | QA Mentor publishes no hourly range on its site or on its directory profiles, so the first number arrives inside a proposal |
| A start time or a commitment floor stated in public | Service page or homepage | QA Madness publishes a one-to-three day start for a dedicated team, and TestDevLab states no minimum commitment and no lock-in. Both are claims a contract can be held to |
No signal decides anything alone. What they do is set the agenda for the first call, so the hour goes on fields the site left open rather than on a capability walkthrough.
Long list, short list, pilot
The long list
Twelve to twenty names is the working size, drawn from directory categories filtered by region and project threshold, and from reference lists of teams shipping a comparable product. Record five fields per candidate: project threshold, published hourly band, certificates with or without numbers, review counts by platform, stated start time. Four of the five sit on directory profiles, so the list takes an afternoon. Expect the start-time column to stay mostly empty: six of the twelve companies here publish one, and a blank cell is itself a finding.
The short list
Three to five names. Apply the two disqualifying requirements first, then the scorecard, then the constraint people skip: how much of the work you still own if the engagement ends in month four. Rank the survivors on completeness of public evidence before capability, because a gap in the evidence becomes a question, a question becomes a call, and the calls stretch procurement from two weeks into two months.
The pilot
Published pilot formats differ enough that the shape of a first order is usually the supplier's rather than yours. DeviQA lists a free proof of concept ahead of any contract. QualityLogic lists pilot programmes alongside its managed and staff augmentation models. TestDevLab lists a one-time engagement from one week upward for audits, pre-launch checks and gap analysis. QA Wolf sells a self-serve platform with no minimum contract and no per-seat charge.
Pick the format that leaves you an artefact. An audit leaves a written finding, a pilot suite leaves test code, a pre-launch check leaves a defect list against a build you shipped. Scope it to one product area, fix the end date, and agree the price before the kick-off call.
Exit criteria, written before the pilot starts
A pilot without a written exit becomes a trial period that never ends. Four gates are enough, each measurable on the day the pilot closes.
- A delivery number against a baseline. Regression run time, automated coverage as a share of the suite, or defects found before release, each with its week-one value recorded beside it.
- Environment readiness. The date the team first ran a test in your environment, counted from the start date. Slippage here predicts the engagement better than any capability answer.
- Handover artefacts. Test cases, automation code and test data in a repository under your organisation, not exported on request from the supplier's tooling.
- The people question. The engineers who ran the pilot are the engineers named in the follow-on order, or the difference is explained in writing.
Convert when three gates are met and the fourth has a dated plan. Where two fail, run the same pilot with the runner-up: a second pilot costs less than a year of the wrong engagement.
When two candidates come out equal
Ties are common, because the fields that separate suppliers are the fields most of them leave blank. This ranking resolves an equal total on the score for the highest-weighted dimension, then the next, in weight order. Apply the same logic with your own weights: the dimension holding veto power settles it, not the total.
If the tie survives, prefer the candidate whose evidence you can confirm without the supplier's help, then prefer the shorter exit. The two shapes on offer are not reversible on the same timescale. a1qa publishes a Build-Operate-Transfer model in which a dedicated team converts into resources the client owns. Kualitatem states contract durations running from six months to five years. Where both candidates still hold, split the pilot across two product areas and let the delivery numbers decide.
The first thirty days
Onboarding fails on access, not on skill. Before day one, prepare the account list, environment credentials, the test data set, repository permissions and one named person who answers questions within two hours. Where a start time is published at all, the figures run from a few days to about two weeks, and the spread describes how much the supplier expects the client to have ready. QASource publishes a five-day average ramp; Kualitatem publishes a one-week onboarding.
Week one belongs to the baseline: current regression run time, current coverage, defect counts by severity, and the release cadence the team must fit. Week two belongs to the first delivered artefact, however small. By day thirty the monthly report should carry the three numbers agreed at the pilot stage with their week-one values beside them. A report that changes its metrics between month one and month three cannot show whether anything improved.
Common questions
How do you choose the right QA outsourcing partner?
Start with the two requirements a candidate has to meet, and write beside each the public artefact that settles it: a register entry, a published rate, a stated start time, a case carrying a figure. Build a long list of twelve to twenty from directory profiles, cut it to three to five, then score the survivors on the same ten rows. Finish with a bounded pilot that has written exit gates. Filtering on public evidence keeps calls for the questions public pages cannot answer.
How do you evaluate a QA outsourcing company before signing a contract?
Evaluate in three passes. Pass one is public material: certificates with numbers, review counts by platform, the project threshold, a stated start time, cases with named clients. Pass two is written answers to what the pages leave open: certificate scope, overlap hours, notice period, ownership of test code, what the rate excludes. Pass three is a paid pilot with a deliverable and a fixed end date. A candidate that clears the first two passes but resists a bounded first order has answered the larger question anyway.
How do you outsource QA testing for the first time?
Begin with a narrow scope you already understand: one product area, one platform, one release. Write the two disqualifying requirements, filter a long list on public evidence, shortlist three to five, and buy an audit or a pilot suite rather than a team. Put the notice period, the asset ownership and the engineer names into that first order. Baseline your own numbers before the team starts, because without week-one figures nothing later can be measured.
How do you onboard an outsourced QA team?
Onboarding is an access problem before it is a testing problem. Prepare accounts, environment credentials, a test data set, repository permissions and one named person who answers questions within two hours. Week one produces the baseline: regression run time, coverage share, defect counts by severity, release cadence. Week two produces the first artefact, however small. A published start time assumes the access list is ready on day one, which is the basis QA Madness states its figure on.
Which QA outsourcing best practices are worth following?
Five hold up across engagement sizes. Filter on commercial fields before capability pages, because a published project threshold settles fit faster than a call. Ask for a certificate number and a registrar rather than an image. Buy a bounded pilot before an annual commitment, and write its exit gates first. Keep test code and test data under your own account from day one. Fix three or four release-level metrics and keep them for the life of the contract.
By Erik Smith, Lead Market Analyst & Founder. Figures in this guide come from the same company files as the ranking; the scales are on the methodology page and every source is listed on the sources page.