OpenMP debugs newbies

I am starting to learn OpenMP, run examples (with gcc 4.3) from https://computing.llnl.gov/tutorials/openMP/exercise.html in the cluster. All examples work fine, but I have a few questions:

  • How to find out in which nodes (or kernels of each node) different threads were launched?
  • The case of nodes, what is the average transmission time to the microsecond or nanosecond for sending information and returning it?
  • What are the best tools for debugging OpenMP programs?
  • Best tips for speeding up real programs?
+7
source share
2 answers
  • Typically, your OpenMP program does not know or care about which kernels it runs on. If you have a job management system that can provide the necessary information in the log files. Otherwise, you can probably insert calls into the environment inside your threads and check the value of some environment variable. What you call and how you do it depends on the platform, I will leave everything you need.

  • How should I know (or any other SOer)? It is clear that you need to tell us more about your equipment, about / s, runtime system, etc. Etc. Etc. The best answer to the question is the one you determine from your own dimensions. I am afraid that you are also mistaken in thinking that information is sent all over the computer - in memory variables with shared memory they usually remain in one place (or at least you should think that they remain in one place, reality may to be much merciless but also impossible to recognize) and is not sent or received.

  • Parallel debuggers like TotalView or DDT are probably the best tools. I haven't used the Intel debugger parallel features yet, but they look promising. I will leave this to less funded programmers than me to recommend FOSS options, but they are there.

  • i) Choose the fastest parallel algorithm for your problem. This is not necessarily the fastest sequential algorithm made in parallel.

    ii) Check and measure. You cannot optimize without data, so you need to profile the program and understand where performance bottlenecks are. Do not believe any line advice that "X is faster than Y". Such statements are usually based on very narrow and often outdated cases and have become “truths” in the minds of their promoters. You can almost always find counter examples. This is YOUR code that you want to make faster, there is no substitute for your research.

    iii) Know your compiler inside out. The return rate (measured by improving the speed of the code) while you were compiling settings is much higher than the rate of return from changing the code "manually".

    iv) One of the “truths” that I draw attention to is that compilers are not very good at optimizing the use of the memory hierarchy in current processor architectures. This is one area where modifying the code might be useful, but you won't know about it until you profile your code.

+7
source
  • You may not know that the thread section on different cores is completely processed by the OS. You are talking about nodes, but OpenMP is multi-threaded (rather than multi-processor) parallelization, which allows you to parallelize one machine containing several cores. If you need parallelization on different machines, you need to use a multiprocessor system such as OpenMPI.

  • The order of magnitude of the communication time:

    • huge in the case of communication between the cores inside the same CPU, it can be considered instantaneous
    • ~ 10 GB / s for data exchange between two processors through the motherboard.
    • ~ 100-1000 MB / s for network communication between nodes, depending on the hardware.

    All theoretical speeds should be indicated in the technical specifications of your equipment. You should also do small tests to know what you really will have.

  • For OpenMP, gdb do a good job, even with many threads.

  • I work in extreme physics on a supercomputer, here are our daily goals:
    • use as little communication between threads / processes as possible, in 99% of cases these are messages that kill the execution of parallel tasks.
    • optimally break down tasks, machine loading should be as close as possible to 100% all the time
    • test, tuning, re-checking, re-tuning .... Parallelization is not at all a general “miraculous solution”, it usually takes practical work to be effective.
+1
source

All Articles