You may not know that the thread section on different cores is completely processed by the OS. You are talking about nodes, but OpenMP is multi-threaded (rather than multi-processor) parallelization, which allows you to parallelize one machine containing several cores. If you need parallelization on different machines, you need to use a multiprocessor system such as OpenMPI.
The order of magnitude of the communication time:
- huge in the case of communication between the cores inside the same CPU, it can be considered instantaneous
- ~ 10 GB / s for data exchange between two processors through the motherboard.
- ~ 100-1000 MB / s for network communication between nodes, depending on the hardware.
All theoretical speeds should be indicated in the technical specifications of your equipment. You should also do small tests to know what you really will have.
For OpenMP, gdb do a good job, even with many threads.
- I work in extreme physics on a supercomputer, here are our daily goals:
- use as little communication between threads / processes as possible, in 99% of cases these are messages that kill the execution of parallel tasks.
- optimally break down tasks, machine loading should be as close as possible to 100% all the time
- test, tuning, re-checking, re-tuning .... Parallelization is not at all a general “miraculous solution”, it usually takes practical work to be effective.
Antonin Portelli
source share