tests/block_cloning: try harder to stay on same txg in fallback test #15303

robn · 2023-09-21T00:44:54Z

Description

We've observed this test failing intermittently. When it does, the "same block" check shows that both files have the same content, that is, the file was cloned.

The only way this could have happened is if the open txg moved between the dd and clonefile calls. That's possible because although we set zfs_txg_timeout to be large, that only affects the wait time in the sync thread at the start of a new txg; it doesn't change anything if its currently waiting or working.

So here we just force the txgs to move immediately before, which should get both operations onto the same txg as intented.

Sponsored-By: OpenDrives Inc.
Sponsored-By: Klara Inc.

How Has This Been Tested?

Ran the whole block_cloning suite on kernels 6.4.2, 6.4.15 and on Fedora 37 specifically.

Types of changes

Bug fix (non-breaking change which fixes an issue)
New feature (non-breaking change which adds functionality)
Performance enhancement (non-breaking change which improves efficiency)
Code cleanup (non-breaking change which makes code smaller or more readable)
Breaking change (fix or feature that would cause existing functionality to change)
Library ABI change (libzfs, libzfs_core, libnvpair, libuutil and libzfsbootenv)
Documentation (a change to man pages or other documentation)

Checklist:

My code follows the OpenZFS code style requirements.
I have updated the documentation accordingly.
I have read the contributing document.
I have added tests to cover my changes.
I have run the ZFS Test Suite with this change applied.
All commit messages are properly formatted and contain Signed-off-by.

We've observed this test failing intermittently. When it does, the "same block" check shows that both files have the same content, that is, the file was cloned. The only way this could have happened is if the open txg moved between the dd and clonefile calls. That's possible because although we set zfs_txg_timeout to be large, that only affects the wait time in the sync thread at the start of a new txg; it doesn't change anything if its currently waiting or working. So here we just force the txgs to move immediately before, which should get both operations onto the same txg as intented. Signed-off-by: Rob Norris Rob Norris <[email protected]> Sponsored-By: OpenDrives Inc. Sponsored-By: Klara Inc.

robn · 2023-09-21T00:46:00Z

@behlendorf this is a little bit of guesswork since I couldn't reproduce it myself, and its probably not technically enough because it doesn't actually lockout the sync in any meaningful way, just tries to get the timing right. But its probably not worse than before! If you've got better ideas let me know!

behlendorf

This works for me (and the CI). It's not the first time we've needed to add a zpool sync to force a txg to to be written specifically for a test case. Thanks for running this down.

We've observed this test failing intermittently. When it does, the "same block" check shows that both files have the same content, that is, the file was cloned. The only way this could have happened is if the open txg moved between the dd and clonefile calls. That's possible because although we set zfs_txg_timeout to be large, that only affects the wait time in the sync thread at the start of a new txg; it doesn't change anything if its currently waiting or working. So here we just force the txgs to move immediately before, which should get both operations onto the same txg as intented. Sponsored-By: OpenDrives Inc. Sponsored-By: Klara Inc. Reviewed-by: Brian Behlendorf <[email protected]> Signed-off-by: Rob Norris Rob Norris <[email protected]> Closes openzfs#15303

We've observed this test failing intermittently. When it does, the "same block" check shows that both files have the same content, that is, the file was cloned. The only way this could have happened is if the open txg moved between the dd and clonefile calls. That's possible because although we set zfs_txg_timeout to be large, that only affects the wait time in the sync thread at the start of a new txg; it doesn't change anything if its currently waiting or working. So here we just force the txgs to move immediately before, which should get both operations onto the same txg as intented. Sponsored-By: OpenDrives Inc. Sponsored-By: Klara Inc. Reviewed-by: Brian Behlendorf <[email protected]> Signed-off-by: Rob Norris Rob Norris <[email protected]> Closes #15303

We've observed this test failing intermittently. When it does, the "same block" check shows that both files have the same content, that is, the file was cloned. The only way this could have happened is if the open txg moved between the dd and clonefile calls. That's possible because although we set zfs_txg_timeout to be large, that only affects the wait time in the sync thread at the start of a new txg; it doesn't change anything if its currently waiting or working. So here we just force the txgs to move immediately before, which should get both operations onto the same txg as intented. Sponsored-By: OpenDrives Inc. Sponsored-By: Klara Inc. Reviewed-by: Brian Behlendorf <[email protected]> Signed-off-by: Rob Norris Rob Norris <[email protected]> Closes openzfs#15303

behlendorf approved these changes Sep 22, 2023

View reviewed changes

behlendorf merged commit 2dc89b9 into openzfs:master Sep 22, 2023
18 of 19 checks passed

Provide feedback

Saved searches

Use saved searches to filter your results more quickly

tests/block_cloning: try harder to stay on same txg in fallback test #15303

tests/block_cloning: try harder to stay on same txg in fallback test #15303

robn commented Sep 21, 2023

robn commented Sep 21, 2023

behlendorf left a comment

tests/block_cloning: try harder to stay on same txg in fallback test #15303

tests/block_cloning: try harder to stay on same txg in fallback test #15303

Conversation

robn commented Sep 21, 2023

Description

How Has This Been Tested?

Types of changes

Checklist:

robn commented Sep 21, 2023

behlendorf left a comment

Choose a reason for hiding this comment