Skip to main content

Command Palette

Search for a command to run...

Remove the code before you delete the flag

Updated
•3 min read•View as Markdown
Remove the code before you delete the flag

A feature flag's expiry date is often wired straight to a deletion job. That has the order backwards. Deleting the configuration is the last step of a removal, not the trigger for one.

The reason is the code default. Every OpenFeature evaluation carries a default the application returns when the provider cannot answer, so once the definition is gone, every call site falls back to that default rather than to the variant someone actually chose. If the rollout ended on candidate and the code default is still control, deleting the definition reverts the behavior silently, at the moment the deletion propagates rather than at a deploy, which leaves no release to correlate the change against.

The order that holds:

  1. Choose the permanent variant and hold it long enough to observe normal load, including the traffic that only appears weekly.

  2. Make the behavior unconditional in code. Remove the tests for the obsolete branch and keep the tests for the behavior that survives.

  3. Deploy everywhere. "Everywhere" is the part that bites.

  4. Confirm silence two ways. Code search finds literal keys and wrappers; evaluation telemetry finds the dynamic and rarely executed consumers that static analysis misses.

  5. Only then delete the targeting rules and the provider definition.

Step 3 fails quietly in multi-service systems, because consumers deploy at different rates. An API that ships several times a week can be clean while a batch worker released monthly still evaluates the same key from an older image. Tracking removal as one boolean per flag hides that. Tracking consumer versions individually does not.

This is also why a useful inventory report carries four states rather than two. Approaching expiry and expired-but-active are the two that teams usually build. The other two are the ones that carry the removal signal: configured but never evaluated means the definition can probably go, and evaluated only through default or error paths means something is asking for a flag the provider is not answering, which is a defect rather than a retirement.

Evaluation telemetry is what makes step 4 an observation instead of an assertion. OpenTelemetry's feature-flag semantic conventions define feature_flag.key, feature_flag.provider.name, feature_flag.result.variant and feature_flag.result.reason, and it is the reason field that separates "a consumer chose this variant" from "a consumer got the default because the provider was unreachable". Emit the semantic variant rather than the raw evaluated value, which can be large or carry context that was never meant to leave the service. Those conventions are still development stage, so pin a version in the instrumentation instead of following attribute names as they move.

Edilec's full article on feature flag lifecycle governance covers the rest of the path: classifying a flag before it exists, per-class failure defaults, keeping evaluation context minimal and typed, review proportional to blast radius, and migrating providers in shadow mode so the decision does not change with the vendor.

AI disclosure: This excerpt was generated with AI from Edilec's published article. It describes documented OpenFeature and OpenTelemetry behavior and a stated engineering practice. It reports no client project, incident or measured result.

M

the part that's easy to skip past is step 1, "hold it long enough to observe normal load." we treat that observation window as the actual gate, not just a waiting period -- pull evaluation telemetry for the whole window and confirm every evaluation landed on the variant you think won, not just that nobody complained.

a kill-switch flipped in the provider dashboard during an incident and never flipped back produces exactly the silent-default failure this post describes, except it looks like a clean removal, because the code default happens to match whatever got served that week by accident. the four-state inventory is the right fix for that -- "evaluated only through default or error paths" catches it -- but only if someone's actually looking at states three and four instead of just expired vs still-active.