Two Step Snapshot Delete #54705

original-brownbear · 2020-04-03T09:31:20Z

Improves the state machine for snapshot deletion to remove race conditions around concurrent snapshot create + delete.
Prior to this change the following would happen during a delete:

Delete operation checks repository if it contains the snapshot to delete.
If one is found -> puts delete operation into the cluster state. If none is found, checks the cluster state for whether the snapshot is in-progress and tries to abort it.
Either waits for the aborted snapshot to finish and then executes step 1 again or physically deletes snapshot from repository after adding the delete to the repo worked out.

This comes with at least two possible problems:
a. If a snapshot finished right after checking the repository contents and before checking the cluster-state for an in-progress snapshot, then the delete will fail with a 404 even though it was sent after the snapshot was started.
b. If an abort finishes and the user concurrently initiates another snapshot, then the retry of the delete might happen after the new snapshot creation is put in the cluster state and will fail because now a snapshot is in progress again.

Currently, both of these are just minor annoyances that led to some test-failures that required workarounds. Problem a. is definitely bad UX though. With concurrent snapshot operations incoming, these kinds of issues are more relevant since they prevent clean ordering of operations if deletes and snapshot creates are allowed to be executed in parallel.

This change adds another step to the delete process and changes the way the delete job is serialized in the cluster state.

Now, the delete always begins with a cluster state update. If that update finds an in-progress snapshot then it will abort it right away and already add the delete job for that snapshot to the cluster state. If no in-progress snapshot is found, then the delete is added to the cluster state with unknown repository state id and unknown SnapshotId.
If no in-progress snapshot is found, then the repository data will be loaded after the placeholder delete has been added to the cluster state and the placeholder is updated with the correct repository generation and SnapshotId found in the repository data.
In the case of an abort, we again wait for the aborted snapshot to finish, then execute the delete.
Since there is a delete in the cluster state, no new snapshot can be started concurrently to mess up the delete operation.
Also, if no in-progress snapshot is found initially then the addition of the delete operation to the cluster state prevents another snapshot from being started concurrently and we get clean ordering of operations in this case as well.

=> Only checking the cluster state contents for in-progress operations during cluster state update jobs and putting a place holder for the delete similar to the snapshot create INIT stage resolves all the above races.
=> The serialization change to the snapshot in progress entries sets up the basis for deleting multiple snapshots at once for #49800 (we can, without another CS serialization change make it so that the name in the delete placeholder is resolved to multiple snapshot ids which is all that's left to make #49800 behave nicely .. currently that PR is blocked on the fact that the races above make multi-delete behave strangely when patterns match in-progress and existing snapshots etc.)
=> The clean ordering of operations from this change gets us another big step closer to having safe concurrent repository operations

WIP: need to clean up the code a little but would like some CI runs

…delete

…ep-snapshot-delete

…delete

elasticmachine · 2020-04-03T09:31:23Z

Pinging @elastic/es-distributed (:Distributed/Snapshot/Restore)

We only have very indirect coverage of master failovers during snaphot delete at the moment. This comment adds a direct test of this scenario and also an assertion that makes sure we are not leaking any snapshot completion listeners in the snapshots service in this scenario. This gives us better coverage of scenarios like elastic#54256 and makes the diff to the upcoming more consistent snapshot delete implementation in elastic#54705 smaller.

…delete

* Add Snapshot Resiliency Test for Master Failover during Delete We only have very indirect coverage of master failovers during snaphot delete at the moment. This comment adds a direct test of this scenario and also an assertion that makes sure we are not leaking any snapshot completion listeners in the snapshots service in this scenario. This gives us better coverage of scenarios like #54256 and makes the diff to the upcoming more consistent snapshot delete implementation in #54705 smaller.

…delete

Snapshot deletes should first check the cluster state for an in-progress snapshot and try to abort it before checking the repository contents. This allows for atomically checking and aborting a snapshot in the same cluster state update, removing all possible races where a snapshot that is in-progress could not be found if it finishes between checking the repository contents and the cluster state. Also removes confusing races, where checking the cluster state off of the cluster state thread finds an in-progress snapshot that is then not found in the cluster state update to abort it. Finally, the logic to use the repository generation of the in-progress snapshot + 1 was error prone because it would always fail the delete when the repository had a pending generation different from its safe generation when a snapshot started (leading to the snapshot finalizing at a higher generation). These issues (particularly that last point) can easily be reproduced by running `SLMSnapshotBlockingIntegTests` in a loop with current `master` (see elastic#54766). The snapshot resiliency test for concurrent snapshot creation and deletion was made to more aggressively start the delete operation so that the above races would become visible. Previously, the fact that deletes would never coincide with initializing snapshots resulted in a number of the above races not reproducing. This PR is the most consistent I could get snapshot deletes without changes to the state machine. The fact that aborted deletes will not put the delete operation in the cluster state before waiting for the snapshot to abort still allows for some possible (though practically very unlikely) races. These will be fixed by a state-machine change in upcoming work in elastic#54705 (which will have a much simpler and clearer diff after this change). Closes elastic#54766

Snapshot deletes should first check the cluster state for an in-progress snapshot and try to abort it before checking the repository contents. This allows for atomically checking and aborting a snapshot in the same cluster state update, removing all possible races where a snapshot that is in-progress could not be found if it finishes between checking the repository contents and the cluster state. Also removes confusing races, where checking the cluster state off of the cluster state thread finds an in-progress snapshot that is then not found in the cluster state update to abort it. Finally, the logic to use the repository generation of the in-progress snapshot + 1 was error prone because it would always fail the delete when the repository had a pending generation different from its safe generation when a snapshot started (leading to the snapshot finalizing at a higher generation). These issues (particularly that last point) can easily be reproduced by running `SLMSnapshotBlockingIntegTests` in a loop with current `master` (see #54766). The snapshot resiliency test for concurrent snapshot creation and deletion was made to more aggressively start the delete operation so that the above races would become visible. Previously, the fact that deletes would never coincide with initializing snapshots resulted in a number of the above races not reproducing. This PR is the most consistent I could get snapshot deletes without changes to the state machine. The fact that aborted deletes will not put the delete operation in the cluster state before waiting for the snapshot to abort still allows for some possible (though practically very unlikely) races. These will be fixed by a state-machine change in upcoming work in #54705 (which will have a much simpler and clearer diff after this change). Closes #54766

original-brownbear · 2020-04-20T10:43:53Z

Will reopen this in a cleaner format shortly

…ic#54866) * Add Snapshot Resiliency Test for Master Failover during Delete We only have very indirect coverage of master failovers during snaphot delete at the moment. This comment adds a direct test of this scenario and also an assertion that makes sure we are not leaking any snapshot completion listeners in the snapshots service in this scenario. This gives us better coverage of scenarios like elastic#54256 and makes the diff to the upcoming more consistent snapshot delete implementation in elastic#54705 smaller.

… (#55456) * Add Snapshot Resiliency Test for Master Failover during Delete We only have very indirect coverage of master failovers during snaphot delete at the moment. This comment adds a direct test of this scenario and also an assertion that makes sure we are not leaking any snapshot completion listeners in the snapshots service in this scenario. This gives us better coverage of scenarios like #54256 and makes the diff to the upcoming more consistent snapshot delete implementation in #54705 smaller.

original-brownbear added 24 commits March 31, 2020 15:39

new style

7559975

bck

23409a1

Merge remote-tracking branch 'elastic/master' into two-step-snapshot-…

0bf5fc2

…delete

bck

f9fec4c

Merge branch 'master' of github.com:elastic/elasticsearch into two-st…

ce06e5e

…ep-snapshot-delete

Merge remote-tracking branch 'elastic/master' into two-step-snapshot-…

1d9689a

…delete

bck

319e9c3

Merge remote-tracking branch 'elastic/master' into two-step-snapshot-…

e41e5c2

…delete

Merge remote-tracking branch 'elastic/master' into two-step-snapshot-…

c5606b1

…delete

fixes

94a3020

fixes

0317f47

Merge remote-tracking branch 'elastic/master' into two-step-snapshot-…

426fad1

…delete

bck

22c21bf

bck

20095a3

Merge remote-tracking branch 'elastic/master' into two-step-snapshot-…

4509157

…delete

all fixed

faf1efa

works nicely

5b85833

Merge remote-tracking branch 'elastic/master' into two-step-snapshot-…

25a91b3

…delete

works

b77bcc8

works

2a32435

Merge remote-tracking branch 'elastic/master' into two-step-snapshot-…

37c4bf1

…delete

works

f471115

Merge remote-tracking branch 'elastic/master' into two-step-snapshot-…

e648cc2

…delete

Merge remote-tracking branch 'elastic/master' into two-step-snapshot-…

feb4e85

…delete

original-brownbear added WIP :Distributed Coordination/Snapshot/Restore Anything directly related to the `_snapshot/*` APIs labels Apr 3, 2020

original-brownbear added 3 commits April 3, 2020 12:48

better assertions

6168bb4

nicer

e6c7173

add comment

288e533

original-brownbear added 4 commits April 7, 2020 10:10

shorter

8acdf63

shorter

986ff1c

shorter diff

984e333

shorter

1bc81ac

original-brownbear mentioned this pull request Apr 7, 2020

Add Snapshot Resiliency Test for Master Failover during Delete #54866

Merged

original-brownbear added 7 commits April 7, 2020 18:44

Merge remote-tracking branch 'elastic/master' into two-step-snapshot-…

67095dd

…delete

nicer diff

82ee958

smaller diff

1f784cb

Merge remote-tracking branch 'elastic/master' into two-step-snapshot-…

a458ba4

…delete

Merge remote-tracking branch 'elastic/master' into two-step-snapshot-…

9c34966

…delete

Merge remote-tracking branch 'elastic/master' into two-step-snapshot-…

18c433c

…delete

fixes

ee01d15

original-brownbear added 4 commits April 9, 2020 11:22

Merge remote-tracking branch 'elastic/master' into two-step-snapshot-…

8442176

…delete

Merge remote-tracking branch 'elastic/master' into two-step-snapshot-…

9628908

…delete

drier

a35b24c

shorter

f8bbc31

original-brownbear marked this pull request as draft April 9, 2020 11:36

much simpler

0da4862

original-brownbear mentioned this pull request Apr 15, 2020

Make Snapshot Deletes Less Racy (#54765) #55226

Merged

original-brownbear closed this Apr 20, 2020

original-brownbear mentioned this pull request Apr 20, 2020

Add Snapshot Resiliency Test for Master Failover during Delete (#54866) #55456

Merged

Provide feedback

Saved searches

Use saved searches to filter your results more quickly

Two Step Snapshot Delete #54705

Two Step Snapshot Delete #54705

original-brownbear commented Apr 3, 2020

elasticmachine commented Apr 3, 2020

original-brownbear commented Apr 20, 2020

Two Step Snapshot Delete #54705

Two Step Snapshot Delete #54705

Conversation

original-brownbear commented Apr 3, 2020

elasticmachine commented Apr 3, 2020

original-brownbear commented Apr 20, 2020