Data Migration When Replacing a System: Cleanup, Field Mapping, Trial Runs, and Cutover

Data Migration When Replacing a System: Cleanup, Field Mapping, Trial Runs, and Cutover | NETVANA Software Insights article cover

During the evaluation stage of a system replacement, the discussion is mostly about whether the features are sufficient, whether the interface is comfortable, and how the cost is worked out. It is only when the switch actually happens that the part most likely to go wrong makes itself known: getting the old data across.

A batch of customer records goes missing after the list is moved, order totals no longer reconcile, one customer turns into three records. Once these surface after cutover they are draining to unpick, and they hit your team’s confidence in the new system immediately. What follows breaks migration into four stages — inventory and cleanup, mapping rules, trial runs and validation, and the cutover — and sets out what each involves and what the client side owns.

Most of the risk in a system replacement sits in the data

New functionality can be tested: whether a button works, whether a flow runs end to end — one pass tells you. Data is different. Whether it is right usually only becomes apparent in real use, and the person who notices is typically a customer or a frontline colleague.

Data problems are rarely a case of “it would not move across,” but of “it moved across and no longer means the same thing.” A field called Notes in the old system actually holds a mixture of customer requirements, sales reminders, and the residue of a campaign run years ago. The new system moves it untouched into a field also called Notes. Technically a complete success; in practice nobody can read it and nobody can search it.

Data migration is therefore not a technical action but an exercise in taking inventory and making decisions. The vendor executes; only the client can answer what counts as correct data. Recognizing that is what makes the schedule and the staffing estimate realistic. If you are still weighing whether to patch, replace, or rewrite, read Legacy System Modernization first.

Stage one: inventory and cleanup

The point of an inventory is to lay out everything the system holds. The questions to answer:

  • Which categories exist (customers, orders, products, inventory, invoices, transaction records, file attachments).
  • Roughly how many records sit in each, and how far back the earliest ones go.
  • Where each category currently lives — the main system, a spreadsheet, a shared folder, or one colleague’s laptop.
  • Who owns each category, and who understands its rules best.

The thing most often missed is data that lives outside the system. In a lot of companies critical information sits in spreadsheets: the customer tiering a salesperson keeps privately, accounting’s reconciliation sheet, the warehouse’s goods-in and goods-out log. None of it is in the old system, so if nobody asks during the inventory it only surfaces after cutover, to be handled separately then.

The inventory also has to cover attachments: scanned contracts, product images, customer uploads. These files are usually numerous, stored inconsistently, and linked to records through a relationship that has to be mapped on its own — effort that is routinely underestimated.

Cleanup follows the inventory. A system that has run for several years will have accumulated problems. Five kinds are common:

  • Duplicates. The same customer becomes several records because a phone number was mistyped or a name carries an extra space.
  • Gaps. Fields that were not mandatory in the early days are empty for the first few years.
  • Inconsistent formats. Some phone numbers carry an area code and some do not; some dates use the ROC calendar (民國) and some the Gregorian; addresses are written every way imaginable.
  • Implausible values. Test records nobody deleted, negative amounts, dates set in the future.
  • Semantic drift. The same field used to store different things in different periods.

The hard part is not technical, it is deciding who sets the rules. When one customer exists as three records, which one they merge into and whose address survives is a business judgment. The division of labor that works: the vendor produces the list of problems, the client assigns a colleague who knows the business to define how each is handled, and the vendor applies those rules in bulk.

The most common misstep is to schedule cleanup after the migration, on the theory that you move everything first and tidy up gradually. That almost never happens: once the new system is live everyone is busy adapting and nobody goes back, so the old system’s problems become permanent features of the new one.

Should all the historical data move?

This is the question asked most often, and the one most often answered wrongly. It also belongs in the inventory stage, because the answer sets the workload for all three stages that follow. The default is not “move everything,” it is “decide by use.”

Sort the data into three groups:

  • Must move. Anything used in daily operations, anything regulation or tax rules require you to keep, anything customers look up.
  • Could move. Historical records consulted occasionally, where moving them helps but leaving them behind still works.
  • Need not move. Data nobody has looked at in years, expired test records, old records unrelated to how the business now operates.

The “regulation or tax rules require you to keep” part of the first group deserves a specific check: retention periods follow tax rules and the regulations covering your particular industry, and they are not the same everywhere. Confirm yours with your accountant or the relevant authority before deciding — do not infer it from the old system’s default setting or from what the previous person in the role happened to do. The statutory periods and legal references for common records — accounting vouchers, books, attendance records, and payroll — are listed under “Retention periods” in Electronic Forms and Approval Workflow Systems.

The third group is usually exported to files for reference, or left in a read-only copy of the old system for a period. Either beats forcing it into the new system: moving it adds migration effort and compromises the new system’s data quality from day one.

A useful test: ask the colleague responsible how far back they have actually looked over the past year. The answer is usually more recent than people assume. What genuinely needs long-term retention is more often a regulatory requirement than a working one, and retention for reference handles that. The logic for moving member and transaction data is discussed further in Building a Membership and CRM System.

Stage two: mapping rules, or how an old field becomes a new one

This step gives every field in the old system an explicit destination in the new one and records the conversion rule. Four situations come up.

A direct one-to-one move. The old customer name maps to the new customer name — the simplest case.

A conversion is needed. The old system used codes 1, 2, and 3 for order status and the new one uses words; or an address was stored as one continuous string and the new system splits it into city, district, and street. Spell out the conversion logic, including what happens to values that match nothing.

A split or a merge is needed. One Notes field has to be broken into several meaningful fields, or several old fields collapse into one. This usually needs human judgment, so at high volume it is worth assessing whether it pays for itself.

Fields the new system lacks, or fields it adds. For an old field the new system does not need, decide whether it is discarded or exported for reference. For a new field the old system never had, decide between a default value and leaving it empty.

Keep the mapping document. It is the only basis you will later have for answering why a record looks the way it does. When somebody questions a figure months afterwards, without it there is nothing to work from but memory.

If you are moving member or customer data, confirm how personal data will be handled and how long it is kept at the same time; How to Write a Website Privacy Policy sets out the principles to check against.

Stage three: trial runs and validation, or confirming it moved correctly

A trial run executes the full migration with real data before the official cutover and loads the result into a test environment for inspection. It cannot be skipped, and it usually happens more than once — the first pass finds the problems, the rules get corrected, a second pass confirms the fix.

Validation has three layers, and all three are worth doing.

First, record counts. How many customers and orders the old system holds, and whether the numbers match after the migration. A count that does not line up is the easiest problem to spot and the first to deal with.

Second, totals and aggregates. Compare a handful of key figures: the order total for a given month, one customer’s accumulated spend, the current total stock on hand. This layer catches the cases where the record count is right but the contents are wrong.

Third, record-by-record spot checks. Someone who knows the business picks a few dozen representative records and compares them directly — the newest, the oldest, the largest, the ones with unusual statuses, the ones already known to be problematic. The most time-consuming layer, and the one most likely to catch semantic errors an automated comparison cannot see.

Validation has to be performed by the client, not read off the vendor’s report. The vendor can confirm the program followed the rules; whether the rules themselves are right is only visible to someone who understands the business. The arrangement mirrors functional acceptance: the people who will use the data decide whether it is correct, and the comparison is recorded in writing.

Stage four: planning the cutover day

Cutover is the day the risk concentrates on, which makes it worth writing out as a timetable in advance, with an owner and an expected duration against each step. The usual sequence:

  1. Announce the downtime window and notify every colleague affected, plus any customers who need to know.
  2. Stop new data entering the old system, putting it into a read-only state.
  3. Take a complete backup of the old system and confirm the backup can actually be restored.
  4. Execute the live migration.
  5. Validate against the agreed list — counts, totals, spot checks.
  6. Once confirmed, open the new system and leave the old one read-only.

Choosing the timing means steering clear of business peaks and month-end closing, which usually puts cutover on a weekend or a public holiday. Note the trade-off: whether the vendor and the internal team can be reached on a day off has to be agreed beforehand.

Setting the rollback criteria in advance is the item skipped most often. Under what circumstances do you abandon the cutover and revert? Who has the authority to call it? By what time? Without deciding beforehand, everyone hesitates under pressure on the day, which drags the situation somewhere harder to recover from.

If the change also alters the URL structure of anything public-facing, handle the search engine side at the same time; The Website Redesign SEO Checklist covers that part.

The watch period after cutover, and the way back

Completing the cutover is not the end of it. The first few weeks after go-live need a defined watch period, focused on three things.

First, make problems easy to report. Tell colleagues explicitly where to send anything that looks wrong, and name one person to collate it. Frontline staff are usually the first to notice, but without an obvious channel they tend to quietly correct things by hand instead.

Second, keep the old system available for a while. Read-only and searchable, so people can go back and check when they are unsure. Decide and announce how long it stays up, to avoid drifting into permanent parallel running — two systems coexisting long term generate a fresh set of inconsistencies of their own.

Third, schedule a review. A month after go-live, look at which data problems were found, which the migration caused, and which the old system had all along. That record has value for any future change to your systems.

The projects where data migration goes well share one thing, and it is not technical brilliance: the preparation period was given enough time for inventory and cleanup, and the client genuinely assigned someone who understands the business to take part in validation. That contribution cannot be outsourced, and it largely decides whether a system replacement lands properly.

What a replacement costs usually depends on how tangled the data is, not on how many features the new system has. To find out where yours sits, show NETVANA what you are working with: we take stock of your existing data and how it is used before discussing migration approach and schedule. Software services are quoted individually rather than sold as fixed packages, and for what each service includes and delivers, see the software services overview.

Further reading: if what you are replacing is a public-facing website, the overall process is covered in How a Website Redesign Actually Runs. The methods for validating data and accepting functionality are much the same, so see How Software Acceptance Works. If the system being replaced is an internal one, read A Guide to Building Internal Admin Systems. And for sequencing the cutover day itself, see The Website Launch Checklist.

Found this useful? Share it