A regression report often starts with a vague observation: the application worked last week, but the same workflow fails today. Reading every intervening change is slow, and blaming the largest patch is unreliable. Git bisect offers a more disciplined approach by repeatedly choosing revisions between a known working state and a known failing state.
The official Git bisect manual describes the search commands, automation contract, and recovery tools. The important preparation is not the command itself. It is defining a test that classifies the same behavior consistently across the history you intend to search. Without that contract, a fast search can produce a convincing but incorrect answer.
Define one observable regression
Start with a specific input and an expected result. A login error for one account type, an incorrect calculation for a fixed fixture, or a repeatable crash is more useful than a general impression that a release feels unstable. Preserve the smallest practical reproducer and explain what passing means.
Separate the target defect from unrelated failures. A historical version might fail to install because its package registry has changed, yet that says nothing about the calculation you are investigating. Your classification must distinguish the regression from an environment that cannot run the test.
Run the proposed test several times on both endpoints. If a supposedly good revision fails intermittently, stabilize the test before searching. Timing-dependent services, mutable remote data, and shared caches can introduce noise. The goal is evidence about a code revision, not evidence about whichever network request happened to succeed.
Prepare a disposable checkout
Bisection switches the checked-out source repeatedly. Use a dedicated clone or worktree and protect any uncommitted work before starting. Do not use a production checkout or a directory containing irreplaceable generated data. A test runner may also create files that must be cleared between revisions.
Record the runtime, dependency installation method, environment variables, and fixture version used for the search. Keep credentials out of the reproducer. If old source requires a different compiler or runtime, plan the supported historical environment rather than quietly testing everything with today's tools.
Place automation scripts outside the searched checkout when practical. The Git manual recommends this arrangement to avoid interactions between the script, the build, and revisions being switched. Otherwise, checking out an older commit may replace the very test that defines your investigation.
Establish good and bad endpoints
Use revisions that you have actually tested, not merely releases that users remember as working. In a disposable checkout, the basic sequence is straightforward:
git bisect start
git bisect bad <known-failing-commit>
git bisect good <known-working-commit>
The angle-bracket values are placeholders to replace with your own verified revision identifiers. Git then selects a candidate revision. Run the same regression test and classify that revision with git bisect good or git bisect bad.
The search assumes a meaningful transition between the endpoints. If the defect was introduced, fixed, and reintroduced, an endpoint choice spanning all those events can complicate interpretation. Narrow the interval around the occurrence you are investigating and record the assumption that the classification represents a consistent boundary.

Skip revisions you cannot evaluate
An unbuildable revision should not automatically be marked bad. If the required toolchain is unavailable or an unrelated failure prevents the reproducer from running, use git bisect skip. That tells Git the revision is untestable under your current procedure.
Skipping is honest, but it reduces the evidence available to the search. The manual warns that skipped commits near the transition can prevent Git from identifying one exact first bad commit. A set of candidates is a legitimate result when the history cannot be tested precisely.
Keep a short reason for each skip. An unavailable dependency and a broken build script may require different follow-up work. If the remaining candidates are important, restoring their historical build environment can be more valuable than repeatedly running the same failing installation command.
Automate only a verified test contract
git bisect run repeatedly executes a command and interprets its exit status. According to the manual, zero means good. Statuses from one through 127 mean bad, except that 125 means untestable and skips the revision. Other statuses abort the process.
This distinction needs deliberate handling. A shell's command-not-found status is 127, which bisection can interpret as bad rather than as a setup error. A missing test executable must not silently become evidence that the application regressed. Validate the wrapper and its dependencies before launching automation.
Make the wrapper translate results explicitly: a confirmed regression should return a chosen bad status, an unrelated inability to evaluate should return 125, and a fatal investigation error should abort. Test all three paths separately. Do not assume the exit code from an arbitrary build tool already expresses your intended classification.
Keep the environment independent of the revision
Build outputs and dependency caches can contaminate neighboring tests. A successful compilation of one revision does not prove that another revision built cleanly if stale artifacts remain. Use the project's documented clean-build process and isolate fixture state where needed.
At the same time, avoid destructive cleanup instructions copied without understanding their scope. A dedicated checkout makes controlled cleanup safer, but it does not justify removing directories outside the investigation. Describe which generated paths the wrapper resets and why.
If the test modifies a database, queue, or local service, restore a known state between candidates. A previous test's successful write can make the next revision appear correct. An operation that depends on an external service also needs a reliable way to distinguish service failure from application behavior.
Preserve the classification record
Use git bisect log to capture the sequence of decisions. Save the output outside the disposable checkout together with the test contract and environment notes. This record explains why the reported boundary was reached.
If you classify a revision incorrectly, the manual documents git bisect replay for rebuilding a corrected investigation from a saved log. Preserve the original record before editing it, remove the incorrect classification deliberately, and replay the corrected version after resetting the session.
Do not turn a bisection log into a substitute for the reproducer. The same good and bad labels mean little if nobody knows what was tested. A useful handoff contains both the search record and enough context for another maintainer to reproduce the important endpoint results.

Validate the candidate before assigning cause
When Git reports the first bad commit, test that revision and its relevant preceding state again under the same conditions. Inspect the change and confirm that the observed behavior matches the original report. A boundary identifies where the test result changed; it does not automatically explain the mechanism.
Consider changes to dependency declarations, generated assets, build configuration, and test fixtures alongside application code. A small configuration change can alter runtime behavior more than a large refactor. Treat the commit as an investigation lead, not an accusation against its author.
Develop a focused fix and add a regression test that captures the intended behavior. Verify the fix on the current supported branch. Reverting the candidate in an old checkout may be informative, but it is not proof that an unreviewed revert is safe in today's release.
Close the session and keep the evidence
Run git bisect reset when the investigation is complete. The documented default restores the original checkout position and removes the bisection state. Check the resulting branch and working tree before returning to ordinary development.
Summarize the tested endpoints, reproducer, skips, candidate boundary, and confirmed explanation. State any remaining uncertainty directly. A result with two plausible candidate commits is better than a false claim of precision.
Git bisect is most useful when it makes an investigation reproducible. Define the behavior first, control the environment, classify honestly, and validate the reported transition. The search then reduces the history you need to inspect without replacing the engineering judgment required to understand and repair the regression.
