Playlist transfer accuracy needs more than a vendor’s headline percentage. Before procurement, freeze your test playlists, define acceptable destination recordings and measure matching separately from writing. Then compare every vendor against the same evidence. A transfer that finds a candidate, chooses the wrong version and writes it successfully should never count as a correct result.
This worksheet targets engineering leads choosing transfer infrastructure for a music service. It provides a sampling matrix, metric definitions, reviewer rules and a hypothetical scorecard. Use it to design your evaluation; it contains no measured vendor rankings.
Playlist transfer accuracy: compare the metrics before the vendors
Start by asking what each percentage counts. A vendor might report candidate selections, completed writes or items already present at the destination. Your acceptance criteria should distinguish those outcomes.
| Metric | Numerator | Denominator | Decision it supports |
|---|---|---|---|
| Match coverage | Source entries with a selected destination candidate | All in-scope source entries | How often does the resolver choose something? |
| Match precision | Selected candidates that reviewers accept | All selected candidates reviewers adjudicate | How often does the resolver choose correctly? |
| Write completion | Selected candidates whose writes succeed | Selected candidates that require a write | Does execution finish? |
| Correct arrival | Source entries with acceptable recordings confirmed at destination | All in-scope source entries | How much of the intended playlist arrives correctly? |
| Correction effort | Active minutes spent correcting results | In-scope source entries, expressed per 1,000 | What operational work remains? |
Treat these as your evaluation definitions. Do not assume a vendor uses identical terminology.
Soundiiz’s playlist transfer tools comparison, checked September 15, 2026, recommends publishing test playlists, service pairs, regions and settings alongside accuracy percentages. Soundiiz sells transfer software, so this supports a testing practice, not an independent ranking.
Keep the scorecard multidimensional. A weighted total can conceal a failure that should block launch, such as replacing explicit recordings with clean versions against your product requirements.
Build a test corpus around your intended service pairs
Make each source-to-destination direction a separate test cell. If your product imports into one destination, prioritize the source services your launch requires. Do not spend your evaluation budget testing irrelevant directions.
Use the following sampling matrix as a worksheet. These rows describe proposed tests, not vendor capabilities or observed results.
| Test cell | Playlist selection | Conditions to freeze | Main question |
|---|---|---|---|
| Each launch direction | Representative listener playlists | Source snapshot, account regions, settings | Does the ordinary import work? |
| Each launch direction | Live, studio, remix, remaster and clean/explicit variants | Version acceptance policy | Does matching preserve recording intent? |
| Each launch direction | Sparse metadata and ambiguous titles | Original metadata and reference evidence | Does the resolver avoid false matches? |
| Priority directions | Long playlists and repeated entries | Entry order, duplicate policy, destination baseline | Does the final playlist preserve structure? |
| Priority directions | Entries with uncertain destination availability | Account region and availability review date | Can reviewers distinguish absence from matching failure? |
Choose representative playlists through a documented selection process. Keep deliberately difficult playlists in a separate challenge set. Report both sets separately so that extra edge cases do not silently change the headline denominator.
Record both account countries. Spotify’s track relinking documentation, checked September 15, 2026, states that track availability depends on the country in the listener’s profile. Your evaluation therefore needs account context, not just service names.
Freeze playlist membership before running vendors. Give each entry an identity based on the playlist snapshot and its position. Count repeated appearances as separate entries when your product promises to preserve them. Also record unique recordings separately, so repeated popular tracks cannot obscure weak performance elsewhere.
Consult the supported features by music service when scoping a MusicAPI evaluation. Confirm the required reads and writes before placing a service pair in the test matrix.
Separate match coverage, match precision and completed writes
Build one worksheet row per source entry. Use your own evaluation columns rather than inventing API fields:
- Run identifier, source service, destination service and account regions.
- Playlist snapshot identifier, source position and source track identifier.
- Source title, artist, version information and available reference metadata.
- Acceptable destination identifiers and the reviewer’s supporting evidence.
- Vendor-selected destination identifier, or no candidate.
- Reviewer verdict: acceptable, incorrect or unresolved.
- Write required, write outcome and destination verification result.
- Correction action, active correction time and final disposition.
Define the denominator before examining results. Let N represent every in-scope source entry, M every entry with a selected candidate and C every selected candidate that reviewers accept. With complete adjudication, coverage equals M/N and precision equals C/M.
If reviewers inspect only a sample, calculate precision within that sample and report the sampling method. Do not divide sampled correct matches by all candidates. Keep unresolved judgments visible and state how many entries still need adjudication.
Measure writes separately. An acceptable candidate that never reaches the destination represents an execution problem. An incorrect candidate that reaches the destination represents a matching problem. The same aggregate completion percentage can hide either. After evaluating matching quality, use playlist transfer reconciliation to account for each source occurrence through destination writes and verification.
MusicAPI uses ISRC-first matching followed by a resolver trained on confirmed human corrections, as its matching approach describes. That approach addresses candidate selection; your reviewers still need to establish recording correctness for your corpus.
Check vendor counter semantics before importing them into your worksheet. MusicAPI’s transfer service documentation describes matchedItems as newly added items and skippedItems as items the destination already had. Neither counter supplies a human correctness judgment. Its hosted results include match rates, unmatched items and a CSV export for evaluation evidence.
For background on identifiers, use ISRCs, track IDs and cross-platform matching. Keep the procurement worksheet focused on outcomes rather than repeating the matching algorithm’s design.
Review versions and unmatched items with explicit acceptance rules
Write the recording policy before reviewers see vendor names. Start with rules like these, then adapt them to your product:
| Source and destination relationship | Proposed verdict |
|---|---|
| Same recording on another compilation | Accept if the product permits release substitution |
| Studio recording replaced by a live performance | Reject |
| Original recording replaced by a remix or cover | Reject |
| Explicit recording replaced by a clean version | Reject unless the listener explicitly permits it |
| Original master replaced by a remaster | Apply the predeclared remaster policy |
| Insufficient evidence to distinguish versions | Mark unresolved and escalate |
Treat identifiers and metadata as evidence within this policy. Do not make a title match, duration tolerance or resolver confidence score the entire acceptance rule.
Have two reviewers independently assess ambiguous cases without vendor labels. Ask a third reviewer to resolve disagreements using the same evidence packet. Log the initial verdicts, final decision and reason. When a policy change affects a class of recordings, reapply it across every vendor.
Review unmatched entries separately. Label each as an acceptable destination recording found manually, no acceptable recording found after the defined search procedure, or unresolved. Avoid declaring that the destination lacks a recording merely because one search failed.
Keep every unmatched entry in the overall coverage denominator. If you also calculate performance on a verified-available subset, publish that subset’s size and selection method. This prevents availability exclusions from making the principal result look stronger.
Finally, inspect playlist order, repeated entries and extra destination items as separate structural checks. A recording-level accuracy metric cannot show whether the import reconstructed the intended playlist. Connect those checks to your playlist migration tool implementation acceptance tests.
Work through a hypothetical transfer scorecard
The following numbers illustrate the worksheet. They do not describe MusicAPI, another vendor or a real test.
Assume one frozen playlist corpus contains 1,000 in-scope entries. The destination starts empty. Reviewers adjudicate every selected candidate, and each selection requires one logical item write. Count final item outcomes after the agreed retry window, not individual request attempts.
| Outcome | Hypothetical count |
|---|---|
| Source entries | 1,000 |
| Entries with selected candidates | 960 |
| Acceptable selected candidates | 936 |
| Incorrect selected candidates | 24 |
| Entries without candidates | 40 |
| Successful item writes | 950 |
| Failed item writes | 10 |
| Acceptable recordings successfully written | 928 |
| Incorrect recordings successfully written | 22 |
The accounting reconciles: 936 acceptable candidates plus 24 incorrect candidates equals 960 selections. Of those selections, eight acceptable recordings and two incorrect recordings fail to write.
Calculate the scorecard:
- Match coverage: 960 / 1,000 = 96.0%.
- Match precision: 936 / 960 = 97.5%.
- Write completion: 950 / 960 = 98.96%.
- Correct arrival: 928 / 1,000 = 92.8%.
- Wrong-version selection rate: 24 / 960 = 2.5%, assuming all incorrect selections involve wrong versions.
A statement that 95% of entries arrived would count 950 writes and conceal 22 wrong recordings. A 97.5% precision statement would describe candidate quality while omitting unmatched entries and write failures. Both figures need their denominators beside them.
Suppose reviewers spend 75 active minutes correcting results. Report 75 minutes per 1,000 source entries, together with the number they corrected and the number they left unresolved. Do not infer that this effort fixes every problem. Keep annotation time, correction time and unattended transfer duration separate.
Set acceptance thresholds and document evaluation limits
Set thresholds before running the comparison. Tie each threshold to a launch requirement and give it an owner. The following gates illustrate a decision structure, not industry benchmarks:
| Gate | Hypothetical threshold | Example result |
|---|---|---|
| Candidate precision | At least 99% in every required cell | Fail: 97.5% |
| Correct arrival | At least 95% in every required cell | Fail: 92.8% |
| Write completion | At least 99% after the agreed retry window | Fail: 98.96% |
| Forbidden substitutions | Zero observed in the challenge set | Requires version-level evidence |
| Correction effort | At most 30 active minutes per 1,000 entries | Fail: 75 minutes |
Do not round a failing result into a pass. Specify minimum sample sizes, retry windows and the treatment of unresolved cases beside the thresholds. Zero observed mistakes in a small sample does not establish zero future risk.
Use the playlist transfer build vs buy worksheet to connect these accuracy acceptance criteria to engineering costs and maintenance ownership.
Publish results per service direction and region before presenting an aggregate. If you weight cells, explain whether the weights reflect measured traffic or planning assumptions. Keep challenge-set performance visible outside the traffic-weighted average.
Retain the corpus version, run timestamps, settings, raw results, reviewer decisions and destination checks. Repeat evaluations after material configuration changes, and distinguish fresh transfers from reruns into populated destinations.
End the evaluation with an explicit decision: accept the tested scope, reject it, or expand testing. List excluded service pairs, untested regions, unresolved judgments and any missing evidence. A reproducible decision remains useful even when the answer is to postpone procurement.
FAQ
Should vendors receive the test playlists in advance?
Share the acceptance policy and a calibration set so vendors understand your requirements. Reserve a separate holdout set for the scored run. Record any configuration changes after calibration, and give each vendor the same opportunity to prepare.
How should procurement handle a vendor that cannot expose selected candidates?
Record an observability gap. You can review final destination entries and measure correct arrival, but you cannot directly calculate candidate precision or isolate matching from write failures. Request the missing evidence before comparing those metrics with vendors that expose it.
Who should approve the final scorecard?
Assign engineering ownership for reproducibility, product ownership for acceptable substitutions and operations ownership for correction effort. Ask procurement to preserve the agreed scope and evidence requirements in the evaluation record.
Bring your frozen corpus and acceptance rules to the MusicAPI transfer service evaluation. For the hosted flow, arrange account enablement and configure a transfer destination before starting.
