ctxp corrupted in _port_irq_epilogue

ChibiOS public support forum for topics related to the STMicroelectronics STM32 family of micro-controllers.

Moderator: RoccoMarco

AndreR
Posts: 24
Joined: Mon Apr 28, 2014 2:01 pm

ctxp corrupted in _port_irq_epilogue

Post by AndreR »

Better grab a coffee Giovanni, that's a long one...

MCU: STM32F303
ChibiOS: 2.6.3 (FPU enabled, prefetch buffer on)
Board: custom

I have a pretty IRQ intense application running (FOC motor control of 2 PMSM) and experiencing spurious BusFaults after some random time (10s - 5mins). So most probably the most nasty thing to debug. After disabling the CM4 write buffer I was able to transform the imprecise BusFaults to precise ones and isolate the problem. The fault always appears at the same code location, which gives me some hope to find a solution.

TIM1, TIM2, TIM3, TIM4 and TIM8 fire fast IRQs with priorities from 3 to 5, ChibiOS is configured this way:

Code: Select all

#define CORTEX_PRIORITY_SYSTICK 9
#define CORTEX_PRIORITY_SVCALL 8
#define PORT_INT_REQUIRED_STACK 128
#define PORT_IDLE_THREAD_STACK_SIZE 512

None of the timers use any OS functions.
Furthermore I have a SPI DMA running (two to be precise, one DMA for RX, one for TX) at IRQ 12 posting to a Mailbox. Here is the code:

Code: Select all

void DMA1_Channel2_IRQHandler()
{
   CH_IRQ_PROLOGUE();

   DMAClearTCIE(DMA_Channel_RX);
   DMAClearTEIE(DMA_Channel_RX);

   msg_t event = EVENT_RX_DONE;

    if(DMA1->ISR & DMA1_FLAG_TE2)
       event = EVENT_RX_ERROR;

    if(SPI1->SR & SPI_FLAG_CRCERR)
       event = EVENT_RX_ERROR;

    chSysLockFromIsr();
    EventClientSendI(spi1.pEventClient, event);
    chSysUnlockFromIsr();

    CH_IRQ_EPILOGUE();
}

EventClientSendI is part of my custom event handling system, it basically does this:

Code: Select all

msg_t EventClientSendI(EventClient* pClient, uint32_t event)
{
   return chMBPostI(pClient->pEventQueue, event | (pClient->clientId << 16)); // combine event with source ID
}


I created a idle tick hook function to measure the CPU load, it goes to 40% at max.

The BusFault occurs during _port_irq_epilogue of the above DMA handler, to be more precise here:

Code: Select all

(ctxp + 1)->fpccr = (regarm_t)(fpccr = SCB_FPCCR);

The ctxp value is 0x55555555, so that dereferencing goes to nirvana. I know that this is the stack fill pattern, but I checked all stacks and they all had large areas of 0x55555555 still there. All stack checks and debug options are enabled, none of them fires.
What I absolutely don't get is why the code crashes in that line. According to the code ctxp must have been used/dereferenced before, but everything is fine there. Is there some weird situation where a fast IRQ could smash in there?

Furthermore, this problem seems to occur only in debug builds. A release build with -Os runs without a problem for at least 30mins. As the BusFaults occur after random time I would not take for granted that the problem never occurs in release builds, I just did not encounter it yet.
I would really like to figure out the root cause of this problem to make sure it does not occur again, please point me to the things to look after/check.
User avatar
Giovanni
Site Admin
Posts: 14891
Joined: Wed May 27, 2009 8:48 am
Has thanked: 1202 times
Been thanked: 996 times

Re: ctxp corrupted in _port_irq_epilogue

Post by Giovanni »

Hi,

Did you check the exceptions stack too? The fact that it only occurs in debug builds makes me thing it could be an overflow of the exceptions stack. Code compiled without optimizations use much more stack than optimized code. The randomness seems to validate this, the crash happen when planets align and multiple ISRs preemt each other.

Could you try to increase the stack size reserved to exceptions?

Giovanni
AndreR
Posts: 24
Joined: Mon Apr 28, 2014 2:01 pm

Re: ctxp corrupted in _port_irq_epilogue

Post by AndreR »

Is that the one:

Code: Select all

__main_stack_size__     = 0x0400;

or this one?

Code: Select all

#define PORT_INT_REQUIRED_STACK 128
User avatar
Giovanni
Site Admin
Posts: 14891
Joined: Wed May 27, 2009 8:48 am
Has thanked: 1202 times
Been thanked: 996 times

Re: ctxp corrupted in _port_irq_epilogue

Post by Giovanni »

Both potentially.

PORT_INT_REQUIRED_STACK is the maximum stack size taken by an IRQ into the thread stack, 32 is fine with optimizations, 64 is fine without optimizations usually.

__main_stack_size__ is the total size of the exception stack, this has to be calculated as the sum of the worst case stack frame size for each IRQ priority level. Something like:

Worst(ISR_frame_priority_0) + Worst(ISR_frame_priority_1) + ... + Worst(ISR_frame_priority_15)

0x400 is usually plenty but if your application is IRQ-intensive and uses a lot of priority levels then this could be a problem.

Giovanni
AndreR
Posts: 24
Joined: Mon Apr 28, 2014 2:01 pm

Re: ctxp corrupted in _port_irq_epilogue

Post by AndreR »

I think I already tested once with main stack size of 0x800, but I will verify this in the evening.

Is there a way to verify that this is a main stack issue by using the debugger? Where and what do I have to look for?
User avatar
Giovanni
Site Admin
Posts: 14891
Joined: Wed May 27, 2009 8:48 am
Has thanked: 1202 times
Been thanked: 996 times

Re: ctxp corrupted in _port_irq_epilogue

Post by Giovanni »

Main and process stacks are placed at the beginning of RAM, see the map file.

There are two possible problems:
1) Main stack overflowing.
2) Process stack overflowing and overwriting the main stack.

You can check the fillers using the debugger after the crash occurred.

Giovanni
AndreR
Posts: 24
Joined: Mon Apr 28, 2014 2:01 pm

Re: ctxp corrupted in _port_irq_epilogue

Post by AndreR »

Sorry for some more questions, I just want to make sure I understand everything correctly to ease information gathering in the evening:

I have the __main_stack from 0x10000000 to 0x10000400 and the __process_stack from 0x10000400 to 0x10000800 (using CCM).

1. Both stacks start initially at 0x10000400 / 0x10000800 and grow downward to smaller addresses, correct? So if I find the 0x55 pattern gone in the e.g. 0x1000000 area I have a problem. "Overflowing" is somewhat misleading, because addresses decrement here.
2. If the process stack overflows it will smash my main stack and thus lead to problems (e.g. corrupt ctxp) during port_irq_epilogue of exceptions, right? But isn't the process stack only used by main()? I put chThdSleep(TIME_INFINITE) to that task in the end, so it is basically unused, how can it overflow?
3. If the main stack overflows it goes into regions below 0x10000000, how can there be a 0x55555555 value then? That's no mans land and probably not initialized to 0x55.

Sorry for so many questions, but RTOS internals are somewhat black magic and CM4 core documentation is rare.
User avatar
Giovanni
Site Admin
Posts: 14891
Joined: Wed May 27, 2009 8:48 am
Has thanked: 1202 times
Been thanked: 996 times

Re: ctxp corrupted in _port_irq_epilogue

Post by Giovanni »

The RTOS is not related to this, the whole stack handling is done in HW in Cortex-M cores.

Your analysis is correct about stacks and how they would overflow. The 0x55 is just a common value in stack areas, an overflow could cause a wrong unstacking and the 0x55 is picked up along the way, not necessarily below 0x10000000.

Other things to check:
1) All normal ISRs must have prologue/epilogue macros.
2) FAST ISRs must NOT have those macros.
3) Verify that priorities are really what you would expect by inspecting the NVIC registers.
4) Do you have the debug options in chconf.h enabled? those can catch a lot of problems.

About main(), do you have large automatic structures declared in that function? that would cause the stack to extend even if no code is executed inside.

Giovanni
AndreR
Posts: 24
Joined: Mon Apr 28, 2014 2:01 pm

Re: ctxp corrupted in _port_irq_epilogue

Post by AndreR »

Okay I did some tests, here are the results:

Normal IRQ macros: checked
Fast IRQ NO macros: checked
NVIC reg (prios): checked
Debug options in chconf.h: all enabled
main(): no large automatic structures
main stack: fine
process stack: ???

The process stack looks suspicous: between 0x10000400 and 0x10000800 is a large area with 0x55 pattern, so that looks good, but:
There are some modified value in the area 0x10000400 to 0x10000430. I set a watchpoint to 0x10000400 to break on write, you can see a screenshot attached. Any ideas what is goind on there? The write happens after some random time, so it think it has something to do with the later fault.
Attachments
stack corruption.png
stack corruption.png (324.96 KiB) Viewed 6159 times
User avatar
Giovanni
Site Admin
Posts: 14891
Joined: Wed May 27, 2009 8:48 am
Has thanked: 1202 times
Been thanked: 996 times

Re: ctxp corrupted in _port_irq_epilogue

Post by Giovanni »

Very interesting, that area should never be touched by the ISR code, can you try if those writes at 0x10000400 happen with optimizations enabled?

I could really use a test case here.

Giovanni
Post Reply