4 ms·
This is really, really cool and I didn't know the platform was open to the extent that you could install your own upstream Linux and just get going. I am curio
by ComputerGuru 2y ago
This is really, really cool and I didn't know the platform was open to the extent that you could install your own upstream Linux and just get going.
I am curious though what configuration option prevents this from ending up with software switching. I understand the mellanox kernel module was compiled and loaded, but certainly that doesn't mean that anything you do in the network stack gets converted to switch fabric code and uses hardware packet switching. How do you make sure that you don't errantly wind up with poor latency and capped throughput?
But also, short of making and selling your own networked device for whatever reason, what are the real benefits of going this approach? I can see crazy use cases for where you have full control of the network stack (but again, see my point above — how do you guarantee you are not doing this in software?) but for most purposes, especially with qsfp and fiber, how much are you really gaining by doing it on-device? What is the killer use case here?
EDIT
Upon rereading, it seems the switch is hard-coded to be hardware-switched and cannot end up in a situation where you are accidentally using software packet switching in the first place (i.e. it does not just optimize to a hardware packet switching state). But that limits what you can do considerably, to the point that an off-the-rack Juniper or CSCO or whatever probably has more features than you can do here without writing your own code to hook into the mellanox sdk?
- benjojo12 2y ago> But that limits what you can do considerably, to the point that an off-the-rack Juniper or CSCO or whatever probably has more features than you can do here without writing your own code to hook into the mellanox sdk? I mean, I'm not touching any mellanox sdk here, I am using the a very similar stack that someone on a "software router" would use, on a switch that can automatically accelerate it to 800G+ throughputs, while hitting a 60W power target. You can hit some of those performance/power numbers in vendor hardware like Juniper/Cisco/Arista, however you have to also put up with their software, I (and others in my group of peers) have not had great experiences with vendor software, and in this setup I am able to patch/fix the software on my own terms. If there is a security vuln in one section, I can fix that, and call it a day, I won't be forced to upgrade parts of the system I do not want to. I cannot do this with Juniper/Cisco/Arista always.
- hamandcheese 2y agoIs it obvious when you try to use a config that won't be accelerated? Or is the config silently ignored?
- benjojo12 2y agoIt really depends on how much you know what you are doing, If you stick to: *) IP Routing that would normally "fit" in a vendor switch *) Bridging *) VRFs You will be fine If you try and do some weird stuff then it's best to check with "ip route" to see if it was actually installed into hardware or not, but I would simply not do anything weird on such hardware
- karma_pharmer 2y agoYes. The switchdev "sw1p[0-9]+" ports are special; the any data the software kernel injects to them is discarded and they never emit packets to the kernel. They exist only to allow you to use `ip bridge` and `ip route` on them. So if you accidentally configure software switching on these ports no data will flow -- it will be totally obvious. You might get "no packets" by accident but you will never get "software switching" by accident. If you really want software switching you have to use the management port (there's only one or two of these) whose name is "eth0" or "eth1" or something like that. So avoiding "accidental software switching" is really easy -- if you're typing "eth" you're doing it wrong. You can even explicitly delete this interface if you don't need the CPU to be able to snoop/inject traffic to/from the switch ports.
- RiverCrochet 2y agoTo configure Linux to do software switching, you need bridge interfaces and NICs to be made part of the bridge with the appropriate `ip` commands (or `brctl` if you're still using that). This is common with small home routers - the WLAN and wired LAN ports (all typically appearing as one NIC) will be made part of a bridge `br0`. The four LAN ports aren't typically exposed as separate NICs so there is hardware switching going on there (some devices do let you split them out because they are VLANed internally though). If `ip link show type bridges` doesn't show any bridges then you aren't software switching unless your drivers are lying to you.
- wmf 2y agoNone of that applies to switchdev; it's a somewhat different world than normal Linux networking.
- karma_pharmer 2y agoI am curious though what configuration option prevents this from ending up with software switching The answer is "switchdev": https://www.kernel.org/doc/html/latest/networking/switchdev.html https://www.kernel.org/doc/html/latest/networking/switchdev.... The Linux switchdev driver is the awesome magic that says "make hardware-offloaded switching ASICs look just like software switching". It's beautiful and amazing, as you'd expect from Mellanox.