osc rdma hang with multiple windows #2530

markalle · 2016-12-06T23:34:31Z

Using the testcase mt_1sided.c (and its support files 1sided.c mt_1sided_td1.c mt_1sided_td2.c) from this test harness pull request:
https://github.com/open-mpi/ompi-tests/pull/25
I get a hang from the osc rdma component:

% mpicc -o x mt_1sided.c mt_1sided_td1.c mt_1sided_td2.c
% mpirun -host hostA,hostB -mca osc rdma -mca pml ob1 -mca btl openib,self,vader ./x

This was from a vanilla build of openmpi-master-201612022109-0366f3a.

The single threaded "1sided.c" passes. And I think the key difference in the multi-threaded version is that there are two windows active (one for each thread).

1sided.c is a fairly broad test, looping over several synchronization types as well as contiguous and non-contiguous datatypes. I could probably whittle the test to be a bit more targeted for this particular hang if needed.

The text was updated successfully, but these errors were encountered:

jsquyres · 2016-12-07T00:24:30Z

@hjelmn Can you have a look?

rkowalewski · 2016-12-07T10:50:37Z

@markalle Your link url does not work. We have probably the same issue regarding communication progress with multiple windows in our project. Could you please fix your url since I want to verify my assumption.

jsquyres · 2016-12-07T11:21:12Z

@rkowalewski Our ompi-test repository is not public; that's why the link is broken for you.

@markalle Are you amenable to sharing your new tests with @rkowalewski? You could post them in a gist, or attach them here on this issue.

rkowalewski · 2016-12-07T14:38:40Z

@jsquyres @markalle I spent more time with this. Finally, I have pinpointed the bug and created another issue: #2538
Maybe there is some relation between these two issues.

markalle · 2016-12-07T18:47:29Z

@jsquyres @rkowalewski Yeah, since we put it in the ompi-tests it's considered public-ish at least. Even the stuff I've checked in there is a whittled down version of our original tests, as giving out these smaller somewhat whittled down tests is easier to get management approval for.

markalle · 2016-12-07T23:23:07Z

Here, see if I made this gist right:
https://gist.github.com/markalle/56f739c3c907e53960d22d5de2fe4bc9
It should have 1sided.c, mt_1sided.c, mt_1sided_td1.c and mt_1sided_td2.c.

jsquyres · 2016-12-07T23:24:19Z

Perfect -- thanks @markalle!

hjelmn · 2017-01-31T15:55:15Z

Please test with the latest 2.0.2 release candidate.

markalle · 2017-06-20T18:33:58Z

I guess we've got two issue reports opened for the same testcase, one for osc=rdma (2530) and one for osc=pt2pt (2614). I updated the other bug report, but will paste the same info here:

I built a vanilla OMPI v2.0.x with --enable-mpi-thread-multiple, and my results were:

mpirun -mca osc rdma -host hostA:2 ... : passed
mpirun -mca osc rdma -host hostA:4 ... : passed
mpirun -mca osc rdma -host hostA:1,hostB:1 ... : segv quickly
mpirun -mca osc pt2pt ... : message that "OSC pt2pt component does not support MPI_THREAD_MULTIPLE in this release"

hjelmn · 2018-02-09T18:10:50Z

Finally found the issue. It is not a threading issue per-se. There is an error in the accumulate master function that ends up leaking requests. Eventually things fall apart and it hangs, crashes, etc. I am testing a fix now.

hjelmn · 2018-02-12T22:25:15Z

Testing is looking good. PR coming shortly.

This commit is a large update to the osc/rdma component. Included in this commit: - Add support for using hardware atomics for fetch-and-op and single count accumulate when using the accumulate lock. This will improve the performance of these operations even when not setting the single intrinsic info key. - Rework how large accumulates are done. They now block on the get operation to fix some bugs discovered by an IBM one-sided test. I may roll back some of the changes if the underlying bug in the original design is discovered. There appear to be no real difference (on the hardware this was tested with) in performance so its probably a non-issue. References open-mpi#2530. - Add support for an additional lock-all algorithm: on-demand. The on-demand algorithm will attempt to acquire the peer lock when starting an RMA operation. The lock algorithm default has not changed. The algorithm can be selected by setting the osc_rdma_locking_mode MCA variable. The valid values are two_level and on_demand. - Make use of the btl_flush function if available. This can improve performance with some btls. - When using btl_flush do not keep track of the number of put operations. This reduces the number of atomic operations in the critical path. - Make the window buffers more friendly to multi-threaded applications. This was done by dropping support for multiple buffers per MPI window. I intend to re-add that support once the underlying performance bug under the old buffering scheme is fixed. - Fix a bug in request completion in the accumulate, get, and put paths. This also helps with open-mpi#2530. - General code cleanup and fixes. Signed-off-by: Nathan Hjelm <[email protected]>

This commit is a large update to the osc/rdma component. Included in this commit: - Add support for using hardware atomics for fetch-and-op and single count accumulate when using the accumulate lock. This will improve the performance of these operations even when not setting the single intrinsic info key. - Rework how large accumulates are done. They now block on the get operation to fix some bugs discovered by an IBM one-sided test. I may roll back some of the changes if the underlying bug in the original design is discovered. There appear to be no real difference (on the hardware this was tested with) in performance so its probably a non-issue. References #2530. - Add support for an additional lock-all algorithm: on-demand. The on-demand algorithm will attempt to acquire the peer lock when starting an RMA operation. The lock algorithm default has not changed. The algorithm can be selected by setting the osc_rdma_locking_mode MCA variable. The valid values are two_level and on_demand. - Make use of the btl_flush function if available. This can improve performance with some btls. - When using btl_flush do not keep track of the number of put operations. This reduces the number of atomic operations in the critical path. - Make the window buffers more friendly to multi-threaded applications. This was done by dropping support for multiple buffers per MPI window. I intend to re-add that support once the underlying performance bug under the old buffering scheme is fixed. - Fix a bug in request completion in the accumulate, get, and put paths. This also helps with #2530. - General code cleanup and fixes. Signed-off-by: Nathan Hjelm <[email protected]>

This commit contains the contents of: 45db363 7f4872d to the v3.1.x branch. These commits fix a couple of bugs and improve the threading support (reference open-mpi#2530). To keep the code mostly in sync with master I added code to osc_rdma_types.h to convert between the atomics support on master and v3.1.x. Signed-off-by: Nathan Hjelm <[email protected]>

This commit contains the contents of: 45db363 7f4872d to the v3.1.x branch. These commits fix a couple of bugs and improve the threading support (reference #2530). To keep the code mostly in sync with master I added code to osc_rdma_types.h to convert between the atomics support on master and v3.1.x. Signed-off-by: Nathan Hjelm <[email protected]>

bwbarrett · 2018-06-12T22:10:01Z

@hjelmn are you ok with closing this ticket?

hjelmn · 2018-06-12T22:15:06Z

Yes.

jsquyres added the bug label Dec 7, 2016

jsquyres mentioned this issue Dec 7, 2016

osc pt2pt wrong answer #2505

Closed

rkowalewski mentioned this issue Dec 7, 2016

Also run the CI tests with shared windows disabled dash-project/dash#184

Closed

fuchsto mentioned this issue Dec 7, 2016

Tests hang using OpenMPI 2.0 dash-project/dash#63

Closed

3 tasks

hppritcha assigned hjelmn Jan 4, 2017

hjelmn added the pending label Jan 31, 2017

rkowalewski mentioned this issue Mar 30, 2017

Problems with MPI_Rget dash-project/dash#350

Closed

hjelmn mentioned this issue Mar 15, 2018

osc/rdma: performance improvments and bug fixes #4918

Merged

hjelmn added this to the v3.1.1 milestone May 8, 2018

hjelmn mentioned this issue May 8, 2018

osc/rdma: bring bug/threading fixes into v3.1.x from master #5158

Merged

bwbarrett modified the milestones: v3.1.1, v3.1.2 Jun 12, 2018

hjelmn closed this as completed Jun 12, 2018

Provide feedback

Saved searches

Use saved searches to filter your results more quickly

osc rdma hang with multiple windows #2530

osc rdma hang with multiple windows #2530

markalle commented Dec 6, 2016 •

edited by jsquyres

Loading

jsquyres commented Dec 7, 2016

rkowalewski commented Dec 7, 2016

jsquyres commented Dec 7, 2016

rkowalewski commented Dec 7, 2016 •

edited

Loading

markalle commented Dec 7, 2016

markalle commented Dec 7, 2016

jsquyres commented Dec 7, 2016

hjelmn commented Jan 31, 2017

markalle commented Jun 20, 2017

hjelmn commented Feb 9, 2018

hjelmn commented Feb 12, 2018

bwbarrett commented Jun 12, 2018

hjelmn commented Jun 12, 2018

osc rdma hang with multiple windows #2530

osc rdma hang with multiple windows #2530

Comments

markalle commented Dec 6, 2016 • edited by jsquyres Loading

jsquyres commented Dec 7, 2016

rkowalewski commented Dec 7, 2016

jsquyres commented Dec 7, 2016

rkowalewski commented Dec 7, 2016 • edited Loading

markalle commented Dec 7, 2016

markalle commented Dec 7, 2016

jsquyres commented Dec 7, 2016

hjelmn commented Jan 31, 2017

markalle commented Jun 20, 2017

hjelmn commented Feb 9, 2018

hjelmn commented Feb 12, 2018

bwbarrett commented Jun 12, 2018

hjelmn commented Jun 12, 2018

markalle commented Dec 6, 2016 •

edited by jsquyres

Loading

rkowalewski commented Dec 7, 2016 •

edited

Loading