4 ms·
I implemented a lockless IPC, DIPC - Interprocess Communication and Distributed IPC library in 2000-2002 time frame. The library is used in vxWorks, QNX, Linux
by srcmap 9y ago
I implemented a lockless IPC, DIPC - Interprocess Communication and Distributed IPC library in 2000-2002 time frame.
The library is used in vxWorks, QNX, Linux (both kernel and user space.) and tested heavily on Xeon SMP 1-2GHz CPU at that time.
For memory management, it was implemented in C++ and the message (Memory) Alloc/Free calls were all lockless. One can allocated a memory from kernel and send it as msg to user space and use it there without copy. The limitation is that the max allows messages are statically pre-allocated for the system. The library manages the pre-allocated pools of various sizes with lockless APIs. When the pools runs out, the API returns error.
The only primitive the library depended on was Atomic_Inc/Dec - absolutely no mutex, spin locks in the non-blocking API code path, etc.
The library was designed to support HA (High available ) TCP/IP stack and routing protocols. The APIs was not the complex. The most complicate part was the regress testing code and the diagnostic code. The API was used by team of 100+ developers. I need help them to debug complex IPC/DIPC issues - and prove to them when the error return, it was not the library/API issue.
There was a trace mechanism that track all the messages went thru the system(s). The trace system was also completely non-blocking and lockless, worked both in kernel and user space. All context switches in Linux Kernel was tracked - very similar to FTRACE in kernel today. I hacked the kernel/bios to preserve the DRAM content part of the trace system during warmboot. This way, I can recovered all the messages, last context switched info even for kernel hang, crash issue for postmortem. The trace system used rdtsc for time stamping had 0.5 nano-seconds resolution on 2Ghz Xeon CPU.
The API/Library worked fine. The overall system complexities was a big issue. We can demo TCP/IP, BGP and various routing protocols that did HA Switch over and In Service OS upgrade. But only as demo - and after switch over the state syncing took a long time and reliability of HA TCP/IP and routing protocol was an big issue.
Base on the lesson learned from that project - I went on and coded a different Application level HA system. That system based everything on top of standard Linux API/utilities. That works much better and the HA port of the SW took only 1 dev (me) 3 weeks to code. It supported in service SW (include Linux kernel) upgrade, HW/SW fault detected and automatically switch over. It was deployed in Comcast. Fault detection took < 50 ms and switch over is almost instance. Full state sync was done with wget over management Ethernet interface and always took < 1 second. All incremental states update were done over UDP and took 1 pkt.
The funny Biz/$ outcomes from these two projects:
The first one: startup sold only one HA router for $200k and burned 100M+ of VC $ but sold to Nokia for $400+Mil. I was able to paid off my mortgage with the options from that company - Not FU money but ok outcome.
The 2nd one: The startup sold $40M of products for $60K from that project to Comcast and various Cable companies at 60% margin. That product (with 19+ FPGA) was created/code/developed with ~8 HW/FW/SW engineers (including the 2 founders). The VC got "a seasoned CEO" - managed to raise a few more rounds, kick out the founders (MIT PhD), hired 160+ people and ran the company to the ground.