Back-End & Infrastructure

The Maintenance Tax Your Adoption Plan Is Missing

A framework migration can finish on schedule and still leave a product team with more work than it can afford to maintain. For a legacy ERP backend, I would treat adoption as a recurring operating commitment, not a one-time delivery estimate. The uncomfortable budget item is the time spent keeping old business behavior intact while the new framework, its dependencies, and its owners change.

The launch estimate understates the work because the maintenance clock starts before launch

Replacing a legacy endpoint with ASP.NET Core 10 creates two systems to understand during the transition: the old implementation that still defines expected behavior and the new one that must reproduce it. That overlap costs more than duplicate hosting, because every defect report requires someone to determine whether the fault is new, inherited, or exposed by a changed dependency. A PM who budgets only for implementation and acceptance testing leaves that investigation without an owner.

The shipping checklist (Can this framework ship on time A PM checklist for adoption) can catch release gaps, but I would add a funded patch calendar because a successful release says little about the cost of the next dependency update. Microsoft’s published support term for .NET 10 LTS is three years; that gives the team a planning horizon, not three years without upgrade work. Security fixes, SDK updates, and changes in adjacent packages can still require code changes and regression tests during that period.

Scope that work against an actual inventory. Record the current .NET SDK, ASP.NET Core and EF Core versions, authentication packages, SQL Server driver, hosting image, and the person who approves updates. GitHub Dependabot can propose package changes, but a proposal is not a tested release because it cannot decide whether a posting rule still produces the right ledger entry. OpenTelemetry traces can show where a request failed, but they cannot establish whether a calculated balance was correct for a particular posting date.

I would reserve a recurring maintenance allocation before approving the migration. An initial planning allowance of two engineer-days per month is a value to revise after the first upgrade cycle, not a claim that every ERP needs that amount. Put package triage, regression-test repair, image refreshes, and production verification against that allowance. If the team already spends more time on those activities, use its recorded time instead. If nobody can name who will spend it, the delivery estimate is incomplete because the work will compete with the next feature after launch.

I would not make framework adoption the default remedy for a slow release process, because a new runtime will not remove approval delays or undocumented posting rules. Ask which recurring burden the change is expected to reduce. If the answer is only “development will be faster,” require a measured baseline for a representative change and a plan to repeat the measurement after the migration.

Database compatibility becomes a permanent test obligation unless someone owns the rules

ERP behavior often lives outside the application layer: stored procedures, SQL Server Agent jobs, triggers, and reports may depend on the same tables. Moving an endpoint to EF Core 10 does not transfer ownership of those dependencies, because a successful object mapping says nothing about the order in which a job and a user action update a balance. That is why I would scope compatibility tests by business operation, not by the number of endpoints converted.

The migration dependency map (Legacy ERP Modernization: Identify Dependencies and Scale Backends) needs an operator column as well as a system column, because a documented dependency is still a production risk if nobody knows who investigates it at 02:00. For each operation in scope, record its source of truth, scheduled consumers, reconciliation method, rollback action, and owner. SQL Server Query Store can identify changed query plans, while a reconciliation query checks whether the records produced by those plans are acceptable.

There is a real implementation choice here. EF Core 10 wins when the team needs tracked updates across a well-understood domain model; it costs mapping upkeep, migration review, and attention to generated SQL. Dapper 2.x wins when a stable stored procedure already owns a complex transaction; it costs explicit SQL ownership and more hand-written mapping tests. Neither choice eliminates database maintenance, because both still depend on table definitions and transactional rules that can change independently of application code.

Give the PM a testable slice rather than a promise of “data parity.” Suppose the team’s measured baseline for a representative posting operation is 420 ms at the 95th percentile in its production-like test environment. Compare the migrated path against that same workload and data shape, because a faster empty-database test would conceal a regression. Separately, tune a 30-minute reconciliation window for the pilot and check whether late-running jobs can make a correct posting appear missing. The latency figure describes an example baseline to obtain from your own measurements; the reconciliation window is an operating choice to validate with the people who close the books.

Keep schema changes and application deployment separable where possible. A backward-compatible column addition can usually be deployed before the code that uses it, whereas dropping a column while an old SQL Server Agent job still reads it creates an avoidable outage. The maintenance cost is the period of dual compatibility and the tests that protect it, so put that period on the schedule rather than calling the schema change done when the migration script succeeds.

Background jobs and identity changes create work that endpoint counts conceal

A migration plan organized around HTTP endpoints misses processes that run without a request. A nightly import may be started by SQL Server Agent, a Windows service, or Hangfire 1.8; moving it changes how retries, credentials, and missed runs are handled. Count each scheduled job as an operational contract, because its failure may not be visible until a later reconciliation or user action.

For every job moved, specify what happens after a partial run. “Retry on failure” is insufficient because a retry can duplicate a financial posting unless the operation has an idempotency key or another documented duplicate check. Prometheus can alert on the age of the last successful run, and Grafana can display that measure, but neither tool decides whether a failed run should be resumed or reversed. Set an initial alert threshold of 90 minutes for a job expected hourly, then change it to match the job’s observed completion spread and the time available to intervene.

Authentication deserves the same treatment. An ASP.NET Core service using OpenID Connect and OAuth 2.0 may validate tokens correctly while breaking an older integration that relies on a different claim name or service account. Inventory token issuers, audiences, scopes, and credential-rotation owners before cutting traffic over, because a test login does not represent unattended connections. Record how a rejected token will be distinguished from a failed business authorization in logs; otherwise support staff will spend time escalating incidents to the wrong team.

I would keep the original scheduler for a stable job during the first application cutover if its run history and recovery procedure are dependable, because changing the runtime and the scheduling system together makes failures harder to attribute. That is a choice to maintain an interface for longer, not an argument to preserve every old job indefinitely. Give the temporary arrangement an owner and a removal condition, such as completing two close cycles without an unexplained reconciliation difference.

A maintenance budget is credible only when routine changes can pass a rehearsal

Before committing the full migration, select one operation with a database write, an authenticated caller, and a downstream job. Rehearse a package update and a rollback on that slice. The exercise reveals maintenance effort because people must locate tests, interpret failures, approve a change, and restore service using the same paths they would use later. Track elapsed time and the roles involved rather than reporting only whether the build passed.

In a repository with the .NET SDK installed, this small Bash check provides a repeatable starting point. It builds and tests a supplied solution, then reports vulnerable direct and transitive packages:

set -eu
solution="${1:?pass a solution file}"
dotnet restore "$solution"
dotnet build "$solution" --no-restore -warnaserror
dotnet test "$solution" --no-build
dotnet list "$solution" package --vulnerable --include-transitive

The vulnerability command produces an inventory, not a complete release gate, because a finding still needs severity assessment, an available fix, and regression testing. Add an owner and due date to findings that cannot be patched immediately. For the pilot, use Pact contract tests if another team consumes the API, and run k6 against a representative workload if latency affects the operation’s deadline; both tests earn their place only when someone will maintain their fixtures and thresholds.

Make the decision with a cost comparison the PM can revisit. Price the migration slice, the parallel-running period, the monthly upkeep allowance, and the cost of retaining the current path. A 5% traffic pilot is a tunable rollout setting, not evidence of safety on its own, because low-volume users may never exercise the posting or close process most likely to fail. Pair the pilot with named operations and explicit rollback triggers, such as a reconciliation difference or a sustained increase over the measured latency baseline.

Do not accept “the new framework is easier to maintain” as an unpriced benefit. It may be easier for developers to change an endpoint while being harder for the organization to operate the full transaction, because the database jobs, identity provider, and support procedures still cross the boundary. Approve the broader move only when the rehearsal shows who maintains those crossings and how much of the team’s capacity they consume.

Start by choosing one posting operation and asking its owner to list every job, credential, query, and reconciliation check it touches. Then schedule a patch-and-rollback rehearsal for that slice before assigning a launch date to the rest. The resulting time record gives you a maintenance estimate you can defend—and a reason to narrow the migration if the team cannot afford to keep the new path healthy.