Most rewrite checklists ask about the code. The thing that actually decides how it goes is whether you can run both systems against the same real traffic and diff the results.
The real specification of a legacy system is not in the docs or the tests. It is the behavior it produces in production, bugs included, the ones downstream consumers have quietly built on. Rewrite from the documented intent and you ship something correct and incompatible.
So before I get to the architecture questions, I settle five operational ones.
What is the observable contract: the outputs other systems actually read, not the interfaces you meant to expose.
Which odd behaviors are load-bearing: rounding, ordering, null handling, retry semantics. Some of those are contracts now, whether you intended that or not.
What is the smallest unit you can cut over: one endpoint, one job, one table owner. If the answer is "all of it", you probably have a launch date and not a plan.
How do you go back: not rollback in theory, but a path that still works after the new system has written data.
Who reads the dual-run. Someone has to look at the diffs daily, or shadow traffic just becomes another ignored log stream.
Across Python 2 to 3 migrations and monolith-to-services rewrites of inherited systems, the ones that went quietly were the ones where the old path stayed alive long enough to act as the judge. The ones that hurt were the ones where correctness got argued in review instead of measured against the thing already running.
I write more on legacy modernization and migration strategy here: https://polycratia.com/c/8d8d9fc4
If you've run a cutover recently, I'd like to know what finally convinced you the new path was safe...