Why isn't optimization happening?
I have C / C ++ code that looks like this:
static int function(double *I) {
int n = 0;
// more instructions, loops,
for (int i; ...; ++i)
n += fabs(I[i] > tolerance);
return n;
}
function(I); // return value is not used.
a built-in compiler function, however it does not optimize n
manipulation. I expect the compiler to be able to recognize that the value is never used once rhs. Is there a side effect that prevents optimization?
doesn't seem to matter, i tried Intel and gcc. Aggressive optimization,-O3
thanks
more complete code (complete code is a repetition of such blocks):
280 // function registers
281 double q0 = 0.0;
282 double q1 = 0.0;
283 double q2 = 0.0;
284
285 #if defined (__INTEL_COMPILER)
286 #pragma vector aligned
287 #endif // alignment attribute
288 for (int a = 0; a < int(N); ++a) {
289 q0 += Ix(a,1,0)*Iy(a,0,0)*Iz(a,0,0);
290 q1 += Ix(a,0,0)*Iy(a,1,0)*Iz(a,0,0);
291 q2 += Ix(a,0,0)*Iy(a,0,0)*Iz(a,1,0);
292 }
293 #endif // not SSE
294
295 //contraction coefficients
296 qK0 += q0*C[k+0];
297 qK1 += q1*C[k+0];
298 qK2 += q2*C[k+0];
299
300 Ix += 3*dim2d;
301 Iy += 3*dim2d;
302 Iz += 3*dim2d;
303
304 }
305 Ix = Ix - 3*dim2d*K;
306 Iy = Iy - 3*dim2d*K;
307 Iz = Iz - 3*dim2d*K;
308
309 // normalization, scaling, and storage
310 if(normalize) {
311 I[0] = scale*NORMALIZE[1]*NORMALIZE[0]*(qK0 + I[0]);
312 num += (fabs(I[0]) >= tol);
313 I[1] = scale*NORMALIZE[2]*NORMALIZE[0]*(qK1 + I[1]);
314 num += (fabs(I[1]) >= tol);
315 I[2] = scale*NORMALIZE[3]*NORMALIZE[0]*(qK2 + I[2]);
316 num += (fabs(I[2]) >= tol);
317 }
318 else {
319 I[0] = scale*(qK0 + I[0]);
320 num += (fabs(I[0]) >= tol);
321 I[1] = scale*(qK1 + I[1]);
322 num += (fabs(I[1]) >= tol);
323 I[2] = scale*(qK2 + I[2]);
324 num += (fabs(I[2]) >= tol);
325 }
326
327
328 return num;
My only guess is potentially a floating point exception that introduced side effects
a source to share
It is used in the code n
, first when it initializes it to 0, and then inside the loop on the left side of the function with possible side effects ( fabs
).
Whether or not you actually use a function return doesn't matter, itself is used n
.
Update: I tried this code in MSVC10 and optimized the whole function. Give me a complete example that I could try.
#include <iostream>
#include <math.h>
const int tolerance=10;
static int function(double *I) {
int n = 0;
// more instructions, loops,
for (int i=0; i<5; ++i)
n += fabs((double)(I[i] > tolerance));
return n;
}
int main()
{
double I[]={1,2,3,4,5};
function(I); // return value is not use
}
a source to share
I think the short answer to this question is just that the compiler can do some optimization in theory, it doesn't mean that it will happen. Nothing comes for free. If the compiler is going to optimize n, then someone has to write code to do it.
This sounds like a lot of work for something that is a fancy corner case and a trivial space saving. I mean, how often do people write functions that do complex calculations just to discard the result? Is it worth writing complex optimizations to reclaim 8 byte stack space in such cases?
a source to share
I can't say for sure if this will have an effect, but you might want to look into the GCC attribute pure
and const
( http://gcc.gnu.org/onlinedocs/gcc/Function-Attributes.html ). It basically tells the compiler that the function only works on its input and has no side effects.
Given this additional information, she can determine that the call is unnecessary.
a source to share
Despite the arguments I had in other threads where all compilers are perfect and never skip optimization. Compilers aren't perfect and don't catch optimizations often.
Cool ones:
int fun (int a)
{
switch (a & 3)
{
case 0: return (a + 4);
case 1: return (a + 2);
case 2: return (a);
case 3: return (0);
}
return (1);
}
For the longest time, if you leave this return at the end, you will get an error that the function set a return type but did not return a value. Some compilers will complain about return () at the end of the function, and complain without it.
From what I can tell from gcc vs say llvm, gcc optimizes within a function within a file, where llvm optimizes whatever it feeds. And you can join the bytecode for the entire project into one file and optimize it all in one shot. Currently gcc output outperforms llvm by ten percent or more, which is interesting. give it time.
Perhaps in your case you are using two inputs that are not declared static (const), so the result of n may change. If it is optimized for the way it works, it cannot shrink it. So my guess is that this is a per-function optimization and the calling function does not know what is affecting the dynamic input I received in the system, even if the return value has not been used, you still need to compute function (I) to solve whatever depends on I. I am assuming it is not an infinite loop, ... means there is a constraint imposed? If not here again dynamic not static, function (I) might be the terminating infinite loop function, or it might be there, waiting for the interrupt service routine to change me and push it out of the infinite loop.
a source to share