Hello,
This is an hot topic I wish to discuss.
I am going through tests on the STM32F7xx which uses the new Cortex-M7. The M7 has two caches for instructions and data, this is the problem, there is no intrinsic coherence mechanisms between DMAs and CPU, I don't know if this is a problem with the STM32F7xx or something that all CM7 devices will experience.
DMA reading from memory
By default the cache works in "write back" mode on RAM, this means that writing to RAM not necessarily updates the memory location but just the cache, the RAM can be written later in time.
This means that writing to a DMA buffer does not work (for example the SPI transmit buffer) as expected.
DMA writing in memory
This is even worse, the DMA writes data in RAM but the CPU can read something different if the buffer is cached.
Now I am cosindering several solutions with various pro and cons.
Leave it to the application
The application is responsible to invalidate the cache before/after accessing DMA buffers, there is a CMSIS API for this. Problems: it would be slow and portability would be lost. Also imagine the support nightmare, most people is not able to debug this kind of stuff.
Automatic invalidation by the drivers
Lots of work on drivers but performance would still be impacted.
Do not enable data cache
Bad performance.
Enable "write through" mode for cache
This only works for CPU writing in RAM but not for DMA writing in RAM. It also requires the use of the MPU.
I am considering a mixed approach:
1) Define a "write through" RAM section, the application would take care to put transmission buffers in this area.
2) Define a "no cache" RAM section, the application would take care to put receive buffers in this area.
3) Alternatively to 1 and 2: just create a "no cache" section for all kind of buffers.
The application would have to declare buffers in a special way (we should provide an abstraction macros because the construct is not standard). This solution would require the use of MPU.
Question and more importantly, isn't the RAM supposed to work with zero wait states? if so, why don't just disable cache on the whole RAM? I have to give this a try before anything else. The data cache would still be active on Flash and external RAMs.
Thoughts?
Giovanni
[TALK] Problems with Cortex-M7
- Giovanni
- Site Admin
- Posts: 14891
- Joined: Wed May 27, 2009 8:48 am
- Has thanked: 1202 times
- Been thanked: 996 times
Re: [TALK] Problems with Cortex-M7
I have used the "separate area for DMA transfers" on NuttX, when I had it configured so that STM32F4 CCM is part of normal heap. This model is very annoying to work with, so I would not recommend making it mandatory.
Biggest problems I encoutered:
My thoughts on what to do:
Biggest problems I encoutered:
- Buffers can come from surprising places. E.g. fwrite(foo, buffer, 1234) might either go to an intermediate buffer (no problem), or the filesystem might pass it directly to the SD card driver if it matches the block size (surprise corruption).
- Small buffers become annoying to allocate. For example if you need a simple SPI transfer of 5 bytes, it would be easiest to allocate it on stack but stack probably would be in cached memory.
My thoughts on what to do:
- Test the impact of disabling cache for RAM. Maybe it is only a small slow-down.
- If impact is small (say, less than 20%), disable cache by default and allow application to enable cache and handle invalidation itself.
- If impact is large, implement invalidation in drivers.
- Giovanni
- Site Admin
- Posts: 14891
- Joined: Wed May 27, 2009 8:48 am
- Has thanked: 1202 times
- Been thanked: 996 times
Re: [TALK] Problems with Cortex-M7
Looking at ST's demos, specifically the "SPI with DMA" demo, they simply put the RAM in "write through" mode but this:
1) Does not address the problem fully.
2) Is worring because it is an half assed attempt at a solution to a known problem.
3) The related code contains also a bug, see below.
Giovanni
1) Does not address the problem fully.
2) Is worring because it is an half assed attempt at a solution to a known problem.
3) The related code contains also a bug, see below.
Code: Select all
static void MPU_Config(void)
{
MPU_Region_InitTypeDef MPU_InitStruct;
/* Disable the MPU */
HAL_MPU_Disable();
/* Configure the MPU attributes as WT for SRAM */
MPU_InitStruct.Enable = MPU_REGION_ENABLE;
MPU_InitStruct.BaseAddress = 0x20010000; <<< This is not aligned with the region size.
MPU_InitStruct.Size = MPU_REGION_SIZE_256KB;
MPU_InitStruct.AccessPermission = MPU_REGION_FULL_ACCESS;
MPU_InitStruct.IsBufferable = MPU_ACCESS_NOT_BUFFERABLE;
MPU_InitStruct.IsCacheable = MPU_ACCESS_CACHEABLE;
MPU_InitStruct.IsShareable = MPU_ACCESS_NOT_SHAREABLE;
MPU_InitStruct.Number = MPU_REGION_NUMBER0;
MPU_InitStruct.TypeExtField = MPU_TEX_LEVEL0;
MPU_InitStruct.SubRegionDisable = 0x00;
MPU_InitStruct.DisableExec = MPU_INSTRUCTION_ACCESS_ENABLE;
HAL_MPU_ConfigRegion(&MPU_InitStruct);
/* Enable the MPU */
HAL_MPU_Enable(MPU_PRIVILEGED_DEFAULT);
}
Giovanni
- Giovanni
- Site Admin
- Posts: 14891
- Joined: Wed May 27, 2009 8:48 am
- Has thanked: 1202 times
- Been thanked: 996 times
Re: [TALK] Problems with Cortex-M7
OK, tests done.
First column is cache with Write Back (default), second is with Write Through, third is with RAM entirely not cached (cache IS enabled, RAM is marked as non-cacheable using MPU). The meaning:
1) Write Through does not have much impact, it is within results fluctuation.
2) Disabling cache on RAM has a huge impact, scores are below those of an F4 (which is much more deterministic).
3) It is not true that RAM is accessed with no wait states, strange because on the F4 that is true.
Unless I am doing something wrong, which is also in the realm of possibilities.
Giovanni
Code: Select all
Test WB WT NoCache F4@168MHz
12.1 1341604 1317063 666659 799993
12.2 1142849 1096439 535974 648644
12.3 1142848 1096438 535975 648644
12.4 4453576 3963280 2261760 2559992
12.5 586233 551017 363100 471906
12.6 913173 921891 627683 699995
12.7 271697 267326 137537 208178
12.8 2230240 2209700 971440 1581160
12.9 3668776 3590636 1845160 1873164
12.10 2138786 2112974 1234386 1366044
12.11 3576340 3692292 3517900 3199988
12.12 2979300 2979296 2201824 1856348
First column is cache with Write Back (default), second is with Write Through, third is with RAM entirely not cached (cache IS enabled, RAM is marked as non-cacheable using MPU). The meaning:
1) Write Through does not have much impact, it is within results fluctuation.
2) Disabling cache on RAM has a huge impact, scores are below those of an F4 (which is much more deterministic).
3) It is not true that RAM is accessed with no wait states, strange because on the F4 that is true.
Unless I am doing something wrong, which is also in the realm of possibilities.
Giovanni
Re: [TALK] Problems with Cortex-M7
Giovanni wrote:strange because on the F4 that is true.
I think it is true because F4 has less RAM. As a result:
1) RAM placed closer to the core
2) Has wider bus
Note: that are only my supposition.
Re: [TALK] Problems with Cortex-M7
Perhaps the cache invalidation should be made part of the DMA functions.
So that dmaStreamEnable() invalidates cache for RAM->peripheral DMA. And dmaStreamDisable() invalidates cache for peripheral->RAM DMA.
So that dmaStreamEnable() invalidates cache for RAM->peripheral DMA. And dmaStreamDisable() invalidates cache for peripheral->RAM DMA.
- Giovanni
- Site Admin
- Posts: 14891
- Joined: Wed May 27, 2009 8:48 am
- Has thanked: 1202 times
- Been thanked: 996 times
Re: [TALK] Problems with Cortex-M7
jpa wrote:Perhaps the cache invalidation should be made part of the DMA functions.
So that dmaStreamEnable() invalidates cache for RAM->peripheral DMA. And dmaStreamDisable() invalidates cache for peripheral->RAM DMA.
Good concept, however I think that cache invalidation should occur when the DMA operation ents, before the DMA callback basically. The problem is that not all drivers use the DMA callbacks, those should be handled case by case. Another exception would be those drivers that do continuous operations, like ADC, those would have to invalidate semi-buffers before the callback.
Ethernet is a separate issue too, it uses its own DMA mechanism.
Giovanni
- Giovanni
- Site Admin
- Posts: 14891
- Joined: Wed May 27, 2009 8:48 am
- Has thanked: 1202 times
- Been thanked: 996 times
Re: [TALK] Problems with Cortex-M7
After sleeping over I decided to proceed as follow and see how it goes:
1) The HAL driver activates the Write Through if the DMA is used. RTC, DTCM, ITCM and external RAMs are excluded by default.
2) The DMA driver exports a function dmaBufferInvalidate(addr, size), this does nothing if there is no cache.
3) Drivers will use dmaBufferInvalidate(addr, size) before calling callbacks or where applicable.
4) The application will be able to use dmaBufferInvalidate(addr, size) too, in case of special requirements.
So far I tested it with the ADCv2 driver and now it works (in the application callback, I will move it into the driver).
Giovanni
1) The HAL driver activates the Write Through if the DMA is used. RTC, DTCM, ITCM and external RAMs are excluded by default.
2) The DMA driver exports a function dmaBufferInvalidate(addr, size), this does nothing if there is no cache.
3) Drivers will use dmaBufferInvalidate(addr, size) before calling callbacks or where applicable.
4) The application will be able to use dmaBufferInvalidate(addr, size) too, in case of special requirements.
So far I tested it with the ADCv2 driver and now it works (in the application callback, I will move it into the driver).
Giovanni
- Giovanni
- Site Admin
- Posts: 14891
- Joined: Wed May 27, 2009 8:48 am
- Has thanked: 1202 times
- Been thanked: 996 times
Re: [TALK] Problems with Cortex-M7
After playing with the device my conclusion is that the cache cannot be made entirely transparent to the application.
The main problem is the alignment of the DMA buffers, those must be necessarily aligned to 32 bytes boundary and the thing cannot be hidden to the application. The reason is that invalidating a buffer could also affect adjacent variables unless the buffer address and size are aligned to cache lines sizes (32 bytes).
Being a complete solution not possible then I think it is pointless to implement a partial solution because the developer has to be aware of cache.
Basically, the application will be responsible for cache handling:
1) Flushing of TX buffers and/or Invalidation of RX buffers.
2) Set up of MPU if a write through mode or non-cacheable areas are desired.
3) Disabling dcache entirely.
4) Other handling strategies.
The HAL shall provide mechanisms for simplified cache handling (part of the new DMA driver) and for MPU setup.
Note that the problem is not specific of STM32, I verified other devices with the Cortex-M7 and the cache is handled in the same way (not handled basically).
Giovanni
The main problem is the alignment of the DMA buffers, those must be necessarily aligned to 32 bytes boundary and the thing cannot be hidden to the application. The reason is that invalidating a buffer could also affect adjacent variables unless the buffer address and size are aligned to cache lines sizes (32 bytes).
Being a complete solution not possible then I think it is pointless to implement a partial solution because the developer has to be aware of cache.
Basically, the application will be responsible for cache handling:
1) Flushing of TX buffers and/or Invalidation of RX buffers.
2) Set up of MPU if a write through mode or non-cacheable areas are desired.
3) Disabling dcache entirely.
4) Other handling strategies.
The HAL shall provide mechanisms for simplified cache handling (part of the new DMA driver) and for MPU setup.
Note that the problem is not specific of STM32, I verified other devices with the Cortex-M7 and the cache is handled in the same way (not handled basically).
Giovanni
- Giovanni
- Site Admin
- Posts: 14891
- Joined: Wed May 27, 2009 8:48 am
- Has thanked: 1202 times
- Been thanked: 996 times
Re: [TALK] Problems with Cortex-M7
I added ADC and SPI demos for the STM32F7xx showing DMA buffers management. Not a big deal but still extra burden on the developers.
The good thing is the developer is free to choose his/her own cache management strategy as I mentioned before.
Giovanni
The good thing is the developer is free to choose his/her own cache management strategy as I mentioned before.
Giovanni