Decide which spreadsheet is authoritative for each field before exporting anything, clean the data while it is still in your hands, load in dependency order - people, then jobs, then balances, then history - run parallel against the old sheets, cut over at a period boundary, then lock the spreadsheets read-only. Migration exposes disagreements that were always there.
Answer that field by field, not file by file, because the honest answer is usually different for each. The payroll sheet may hold the correct salary while the HR master holds the correct designation and the attendance file holds the only accurate joining dates. Write a short table listing every field you intend to load, the file it will come from and the person who confirms it. That table becomes the specification for the whole migration and settles arguments later, when two numbers disagree and nobody remembers which source was chosen. Building a single [employee database](/employee-database-software) is the goal, but you cannot merge sources until you have decided which one wins for each column.
Because cleaning inside a live system is slower, riskier and visible to everyone. In a spreadsheet you can sort, filter, compare columns and fix a hundred rows in one operation. In a system each correction is a transaction, possibly requiring an approval, possibly triggering a recalculation, and always leaving a trail that someone will later have to interpret. There is also a sequencing trap: load first and you will be cleaning while people are already transacting, so the data moves underneath you. Fix name formats, duplicate people, impossible dates, blank identifiers and inconsistent salary structures before anything is imported, and validate the file against the target format rather than against your own eye.
Load in the order the system needs things to exist. Organisation structure, locations and legal entities come first, since a person cannot be assigned to a department that has not been created. People come next, with the identifiers everything else will reference. Jobs, grades, managers and pay structures follow, because they attach to a person. Balances and accruals come after that, since they belong to a policy that must already be configured. History comes last, if at all. Loading out of order produces orphaned records and forced re-imports, and a [core HR record](/hris) built on a broken first load is far harder to repair than to redo. Import into a test environment first, every time.
Enough to operate and to answer the questions you can reasonably expect, and no more. Current balances, employment status changes and salary history are usually worth carrying because they get queried. Old attendance detail, superseded policy versions and archived approvals rarely are, and each of them multiplies the cleansing work. A practical split is to load what a person or their manager might look up, and to retain the rest as exported files stored under your retention policy so it remains available without cluttering the system. Decide this early, because history scope quietly drives the size of the whole exercise and it is the easiest thing to over-order at the start.
Reconcile at three levels before anyone is allowed to transact. Counts first: people loaded against people expected, by department and by status, with the differences explained rather than accepted. Totals next: aggregate salary and total leave liability compared against the source. Then individual spot checks, deliberately choosing awkward cases - the person who changed grade mid-year, the one on unpaid leave, the transfer between entities, the rehire. Sign each level off with a name against it. Reconciliation done by the same person who prepared the file catches fewer errors than reconciliation done by someone who did not, which is a good reason to split those roles.
Freeze them rather than delete them. Once the system is live, move every source file to a read-only location with a dated name, tell people plainly that it is historical, and remove edit access from everyone including yourself. Files left editable get updated by someone acting in good faith, and within a couple of cycles you have two competing records again. Announce the change with the same seriousness as the go-live itself, because the sheets are habit and habits outlive announcements. Keep the archive accessible for the retention period you have committed to, since the migration table you built at the start is exactly what you will need if a historical figure is ever questioned.
Because it forces agreement on facts that were previously allowed to differ quietly. Two sheets held two joining dates and nobody had to reconcile them; the system accepts only one. A leave balance calculated one way by HR and another way by the team lead has to become a single number that someone will be paid on. A designation used in a report turns out never to have matched the one on the contract. None of these are created by the migration - they were always there, absorbed by human judgement each time they surfaced. Expect the conversations, schedule time for them, and treat the decisions as policy outcomes rather than data-entry problems.
Pitch N Hire is an applicant tracking system built for recruiters and hiring teams. If this answer described something you want to run properly, the ATS is where it lives.
Free for 1 user · No credit card · Talk to a real hiring expert
Get a personalized walkthrough of Pitch N Hire on your own roles and workflow. No slides, no obligation.
Prefer to talk? Book a demo · Talk to sales · View pricing
Free 1-user plan · No credit card · Talk to a real hiring expert
See your true cost-per-hire and how much Pitch N Hire could save you — our free Recruitment ROI Calculator gives you the numbers in under a minute. No signup required.
Open the free ROI calculatorPrefer a tailored walkthrough on your real roles? Drop your work email:
★ Free 1-user plan · No spam · Talk to a real hiring expert