Why isn't optimization happening?

I have C / C ++ code that looks like this:

static int function(double *I) {
    int n = 0;
    // more instructions, loops,
    for (int i; ...; ++i)
        n += fabs(I[i] > tolerance);
    return n;
}

function(I); // return value is not used.

      

a built-in compiler function, however it does not optimize n

manipulation. I expect the compiler to be able to recognize that the value is never used once rhs. Is there a side effect that prevents optimization?

Compiler

doesn't seem to matter, i tried Intel and gcc. Aggressive optimization,-O3

thanks

more complete code (complete code is a repetition of such blocks):

  280         // function registers
  281         double q0 = 0.0;
  282         double q1 = 0.0;
  283         double q2 = 0.0;
  284
  285 #if defined (__INTEL_COMPILER)
  286 #pragma vector aligned
  287 #endif // alignment attribute
  288         for (int a = 0; a < int(N); ++a) {
  289             q0 += Ix(a,1,0)*Iy(a,0,0)*Iz(a,0,0);
  290             q1 += Ix(a,0,0)*Iy(a,1,0)*Iz(a,0,0);
  291             q2 += Ix(a,0,0)*Iy(a,0,0)*Iz(a,1,0);
  292         }
  293 #endif // not SSE
  294
  295         //contraction coefficients
  296         qK0 += q0*C[k+0];
  297         qK1 += q1*C[k+0];
  298         qK2 += q2*C[k+0];
  299
  300         Ix += 3*dim2d;
  301         Iy += 3*dim2d;
  302         Iz += 3*dim2d;
  303
  304     }
  305     Ix = Ix - 3*dim2d*K;
  306     Iy = Iy - 3*dim2d*K;
  307     Iz = Iz - 3*dim2d*K;
  308
  309     // normalization, scaling, and storage
  310     if(normalize) {
  311         I[0] = scale*NORMALIZE[1]*NORMALIZE[0]*(qK0 + I[0]);
  312         num += (fabs(I[0]) >= tol);
  313         I[1] = scale*NORMALIZE[2]*NORMALIZE[0]*(qK1 + I[1]);
  314         num += (fabs(I[1]) >= tol);
  315         I[2] = scale*NORMALIZE[3]*NORMALIZE[0]*(qK2 + I[2]);
  316         num += (fabs(I[2]) >= tol);
  317     }
  318     else {
  319         I[0] = scale*(qK0 + I[0]);
  320         num += (fabs(I[0]) >= tol);
  321         I[1] = scale*(qK1 + I[1]);
  322         num += (fabs(I[1]) >= tol);
  323         I[2] = scale*(qK2 + I[2]);
  324         num += (fabs(I[2]) >= tol);
  325     }
  326
  327
  328     return num;

      

My only guess is potentially a floating point exception that introduced side effects

+2


a source to share


4 answers


It is used in the code n

, first when it initializes it to 0, and then inside the loop on the left side of the function with possible side effects ( fabs

).

Whether or not you actually use a function return doesn't matter, itself is used n

.



Update: I tried this code in MSVC10 and optimized the whole function. Give me a complete example that I could try.

#include <iostream>
#include <math.h>

const int tolerance=10;

static int function(double *I) {
    int n = 0;
    // more instructions, loops,
    for (int i=0; i<5; ++i)
        n += fabs((double)(I[i] > tolerance));
    return n;
}


int main()
{
    double I[]={1,2,3,4,5};

    function(I); // return value is not use
}

      

+7


a source


I think the short answer to this question is just that the compiler can do some optimization in theory, it doesn't mean that it will happen. Nothing comes for free. If the compiler is going to optimize n, then someone has to write code to do it.



This sounds like a lot of work for something that is a fancy corner case and a trivial space saving. I mean, how often do people write functions that do complex calculations just to discard the result? Is it worth writing complex optimizations to reclaim 8 byte stack space in such cases?

+2


a source


I can't say for sure if this will have an effect, but you might want to look into the GCC attribute pure

and const

( http://gcc.gnu.org/onlinedocs/gcc/Function-Attributes.html ). It basically tells the compiler that the function only works on its input and has no side effects.

Given this additional information, she can determine that the call is unnecessary.

+1


a source


Despite the arguments I had in other threads where all compilers are perfect and never skip optimization. Compilers aren't perfect and don't catch optimizations often.

Cool ones:

int fun (int a)
{
   switch (a & 3)
   {
      case 0: return (a + 4); 
      case 1: return (a + 2);
      case 2: return (a);
      case 3: return (0);
   }
   return (1);
}

For the longest time, if you leave this return at the end, you will get an error that the function set a return type but did not return a value. Some compilers will complain about return () at the end of the function, and complain without it.

From what I can tell from gcc vs say llvm, gcc optimizes within a function within a file, where llvm optimizes whatever it feeds. And you can join the bytecode for the entire project into one file and optimize it all in one shot. Currently gcc output outperforms llvm by ten percent or more, which is interesting. give it time.

Perhaps in your case you are using two inputs that are not declared static (const), so the result of n may change. If it is optimized for the way it works, it cannot shrink it. So my guess is that this is a per-function optimization and the calling function does not know what is affecting the dynamic input I received in the system, even if the return value has not been used, you still need to compute function (I) to solve whatever depends on I. I am assuming it is not an infinite loop, ... means there is a constraint imposed? If not here again dynamic not static, function (I) might be the terminating infinite loop function, or it might be there, waiting for the interrupt service routine to change me and push it out of the infinite loop.

0


a source







All Articles