Need explanations about Floating Point Unit

ChibiOS public support forum for topics related to the STMicroelectronics STM32 family of micro-controllers.

Moderator: RoccoMarco

User avatar
Giovanni
Site Admin
Posts: 14891
Joined: Wed May 27, 2009 8:48 am
Has thanked: 1202 times
Been thanked: 996 times

Re: Need explanations about Floating Point Unit

Post by Giovanni »

I just compiled the IRQ_STORM_FPU demo for the F4 and it generates FPU instructions:

Code: Select all

  56 0000 06EE100A       fmsr   s12, r0
  57 0004 06EE901A       fmsr   s13, r1
  58 0008 00EE102A       fmsr   s0, r2
  59 000c 00EE903A       fmsr   s1, r3
  60 0010 36EE267A       fadds   s14, s12, s13
  61 0014 70EE207A       fadds   s15, s0, s1
  62                    .loc 1 23 0
  63 0018 27EE271A       fmuls   s2, s14, s15
  64 001c 11EE100A       fmrs   r0, s2


Giovanni
pito
Posts: 199
Joined: Sun Nov 06, 2011 3:54 pm

Re: Need explanations about Floating Point Unit

Post by pito »

It might be we simply link the math lib without recompiling it for FPU. I would be happy to see a difference in the trigonometric demo I posted - my naive understanding is the demo w/ FPU on shall be at least 2x faster compared to sw math :)

PS: I can see a diff when doing simple mult/div (8x), the -DUSE_FLOAT=TRUE in Makefile has no effect, the USE_FPU = yes does.
Now, how to recompile the math libs.. The time diff in above trigo test is 111 vs 114, so it seems the math libs are sw double.

The CMSIS-DSP includes arm_cortexM4lf_math.lib (Little endian and Floating Point Unit on Cortex-M4)
pito
Posts: 199
Joined: Sun Nov 06, 2011 3:54 pm

Re: Need explanations about Floating Point Unit

Post by pito »

I've added the path to fpu libs libm.a
C:\ChibiStudio\tools\GNU Tools ARM Embedded\4.7 2013q2\arm-none-eabi\lib\armv7e-m\fpu
C:\ChibiStudio\tools\GNU Tools ARM Embedded\4.7 2013q2\arm-none-eabi\lib\armv7e-m\softfp


and changed:

Code: Select all

q32= asinf(acosf(atanf(tanf(cosf(sinf(q32))))));


With

Code: Select all

-mfloat-abi=hard

I get 31ms (vs 111ms) so ~3.5x speedup.. Good.. :D

With

Code: Select all

-mfloat-abi=softfp

I get 57ms (vs 111ms) so ~2x speedup.. Not so good.. :(
User avatar
Giovanni
Site Admin
Posts: 14891
Joined: Wed May 27, 2009 8:48 am
Has thanked: 1202 times
Been thanked: 996 times

Re: Need explanations about Floating Point Unit

Post by Giovanni »

If this works then it would be a good idea to change defaults, I don't know if there are hidden problems however in using the hard compiled libraries.

Giovanni
pito
Posts: 199
Joined: Sun Nov 06, 2011 3:54 pm

Re: Need explanations about Floating Point Unit

Post by pito »

As a first step I would try to link the original STM hard fpu libm.a if any - I would expect much better results as with the above gcc one.
pito
Posts: 199
Joined: Sun Nov 06, 2011 3:54 pm

Re: Need explanations about Floating Point Unit

Post by pito »

Nope, the hard does not work properly. It seems it just passes the argument through the function - see sinf (for example):

-mfloat-abi=hard:

Code: Select all

 sinf( 1 ) = 1.00000
 sinf( 2 ) = 2.00000
 sinf( 3 ) = 3.00000
 sinf( 4 ) = 4.00000
 sinf( 5 ) = 5.00000
 sinf( 6 ) = 6.00000
 sinf( 7 ) = 7.00000
 sinf( 8 ) = 8.00000
 sinf( 9 ) = 9.00000
 sinf( 10 ) = 10.00000


-mfloat-abi=softfp:

Code: Select all

 sinf( 1 ) = 0.84147
 sinf( 2 ) = 0.90929
 sinf( 3 ) = 0.14112
 sinf( 4 ) = -0.75680
 sinf( 5 ) = -0.95892
 sinf( 6 ) = -0.27941
 sinf( 7 ) = 0.65698
 sinf( 8 ) = 0.98935
 sinf( 9 ) = 0.41211
 sinf( 10 ) = -0.54402

:evil:
User avatar
Giovanni
Site Admin
Posts: 14891
Joined: Wed May 27, 2009 8:48 am
Has thanked: 1202 times
Been thanked: 996 times

Re: Need explanations about Floating Point Unit

Post by Giovanni »

It could be matter of problems in the Makefile, are the options passed to both compiler and linker?

Giovanni
pito
Posts: 199
Joined: Sun Nov 06, 2011 3:54 pm

Re: Need explanations about Floating Point Unit

Post by pito »

Ok, it seems this is the fix:

In rules.mk:

Code: Select all

MCFLAGS   = -mcpu=$(MCU) -mfloat-abi=hard -mfpu=fpv4-sp-d16 -fsingle-precision-constant


in Makefile:

Code: Select all

# End of user defines
##############################################################################

ifeq ($(USE_FPU),yes)
  USE_OPT += -mcpu=cortex-m4 -mfloat-abi=hard -mfpu=fpv4-sp-d16 -fsingle-precision-constant
  DDEFS += -DCORTEX_USE_FPU=TRUE
else
  DDEFS += -DCORTEX_USE_FPU=FALSE
endif


Now the above test demo with hard shows:

Code: Select all

 THIS IS THE START
 Elapsed time 1000x : 9 millis
 Elapsed time : 0.00899 millis
 Result : 9.00001
 THIS IS THE END


The same results with softfp.

Code: Select all

 THIS IS THE START
 Elapsed time 1000x : 9 millis
 Elapsed time : 0.00899 millis
 Result : 9.00001
 THIS IS THE END


The results with FPU off (in Makefile) and with the original rules.mk, single precision ( q32= asinf(acosf(atanf(tanf(cosf(sinf(q32)))))); ):

Code: Select all

 THIS IS THE START
 Elapsed time 1000x : 59 millis
 Elapsed time : 0.05900 millis
 Result : 9.00001
 THIS IS THE END


The results with FPU off (in Makefile) and with the original rules.mk, but double precision ( q32= asin(acos(atan(tan(cos(sin(q32)))))); ) :

Code: Select all

 THIS IS THE START
 Elapsed time 1000x : 114 millis
 Elapsed time : 0.11400 millis
 Result : 9.00000
 THIS IS THE END


So 9 ms against 114ms (13x faster than double precision), and 9ms against 59ms (6.5x faster than single precision) - unbelievable, there still must be an issue somewhere :)
Post Reply