on multiple fronts#
Sun 12 Feb 2023 11:04:08 PM UTC
This week I have been dabbling on multiple fronts.
Running tests in CI
Building kernels and squashfs images in Hydra is a Good Start, but it
doesn't give much confidence that the image is going to do anything
when run. To test this, we can build an image for the "malta" MIPS
variant and then run it in QEMU. To test that it does anything
useful, considering that the device is supposed to be a router, we
need other network devices for it to connect to. And as we're running
it all in Hydra which uses a private network namespace with no
interfaces except loopback, those other network devices need to be
provided as part of the test setup.
As I mentioned last week, we currently use RouterOS running in QEMU to
provide a simulated ISP. The QEMU VMs for RouterOS and our Liminix
device are configured with emulated network devices using QEMU socket networking, and
because of the private network namespace we configure it with
localaddr=127.0.0.1. I'm not sure that that option is documented
anywhere other than the QEMU source
qemu-system-mips
[... a raft of other options....]
-netdev socket,id=access,mcast=230.0.0.1:1234,localaddr=127.0.0.1 \
-device virtio-net-pci,disable-legacy=on,disable-modern=off,netdev=access,m
ac=ba:ad:1d:ea:21:02 \
The nice things about using multicast for this: (1) with regular UDP
sockets we have to connect in one VM and listen in the other,
meaning that we have to start them in a particular order; (2) we can
have more than two VMs attached to the same network.
The other interesting part of this work was scripting RouterOS so that
its VM would boot with a useful configuration instead of with a blank
fresh install. (Admire the irony: in order to build a router operating
system which can be easily managed using version-controlled text
files, I need to force another router operating system which was not
designed with that goal to be managed using a version-controlled text
file). RouterOS on QEMU can be provisioned using the QEMU Guest
Agent - while
I couldn't find a general-purpose host-side utility that communicates
with the guest agent, there is a page on the Mikrotik Wiki with an
example
that I was able to
adapt.
Running builds in CI
I copied and adjusted the configuration for GL.Inet
Mango and
Azure from NixWRT. I
haven't yet tested booting either of them as I don't have any suitable
hardware devices that aren't in use somewhere. This adds to the
existing GL-AR750 support to give us three hardware targets plus QEMU
Packaging cleanups
-
we were amassing a small collection of shell scripts and utilities
for the build machine. For better DX I turned these into derivations
available under liminix.buildEnv
-
found and fixed packaging bugs in the Liminix development TFTP
server caused by Nixpkgs fennel changed. You may reasonably ask why
Liminix even contains a TFTP server - I may have got carried away, but
the USP is that it has an allow-list for client connections and it
follows symlinks. This is desirable for Liminix testing because you
probably don't want to allow access from anything other than the
device you're testing on, and you want to point it at ./result
without messing about copying files into /var/tftp and
what-have-you. Anyway, have now added that as a CI target as well
so we'll know if it breaks again.
-
turned the per-device configuration from an
overlay-plus-a-kconfig-attrset into a module. Given that modules can
set kernel config options (e.g. you can say
kernel.config.SERIAL_OF_PLATFORM = "y") there seemed no point in
having two mechanisms. This means that per-device code can now change
any other part of the config not just the kernel, so also there's no
extra logic needed for setting the boot commandline or default output
format or anything like that. I love removing special cases, it's my
favourite part of programming.
-
rearrange the PHRAM/TFTP boot support to not be obviously broken,
and also to calculate image sizes and offsets instead of leaving that
to the user. No more weird bugs where the kernel stomps on the start
of the root image. Whether it is non-obviously broken I will find
out next week.
Yes, next week I'll be looking at hardware. I have hooked up the
GL-AR750 to a serial console, a remote-controlled power switch, and a
spare ethernet card in my desktop. Next I need to
- try booting the image that Hydra built
- figure out wlan support, which I remember as being gnarly but hopefully my
extensive notes from last time will help with
- connect it to something upstream. I'd like it to run PPPoE (because
that's a typical use case for a home wifi router) and I'd like to hook
it up to the AA.net L2TP service (because that
means I don't need to break broadband for the rest of the household while I'm hacking on it). I am hoping that go-l2tp can bridge the gap
Once I have something I can dogfood then I hope things will settle down
and I can focus on one thing at a time.
Sub-liminix messaging#
Wed 15 Feb 2023 09:23:48 PM UTC
I am restarting/rewriting NixWRT,
he said, a few months ago. This is a short follow-up announcement
to say that
I am very stoked about this. I'm aiming for ~ weekly updates in that
place.
Tunnel Vision#
Sun 19 Feb 2023 11:42:36 AM UTC
This week in Liminix: am happy to report that the rewritten
PHRAM/tftpboot stuff almost worked first time, and the delta between
first and second time was just a simple syntax error. The GL-AR750
boots :-)
Most of this week was spent on tooling and infra to make Actual
Hardware Development simpler.
Trivial Fix, Transmits Packets
I mentioned last week that I'd written a TFTP server that
plays nice with Nixpkgs conventions (doesn't expose all of /,
but does follow symlinks so I can serve files like result/uimage).
(When I say "written" it's really a thin wrapper around someone
else's library). After hooking up to a client, it turned out not to be
100% bug free, so spent some time fixing problems in the way it uses luasocket.
Border Network Gateway
(A tip: performing internet searches for the term "BRAS" - Broadband
Remote Access Server - may not yield the intended results if you don't
add some qualifiers)
Here in the UK, the interface between your home router and the
internet that comes into the house is probably PPPoE (PPP over
Ethernet). Unless it's DHCP ...
-
If you have FTTP (fibre to the premises) there's an Openreach
ONT
somewhere in your house that has an optical port for the incoming
fibre connection and an ethernet port for, in their words, "an
Ethernet cable that runs to your BT Hub". A PPPoE server is running on
this port.
-
If you have VDSL or ADSL, you have a "DSL modem" that performs a
similar role except that the upstream is a RJ11 plug that attaches to
the phone line. Sometimes your ISP send you a combined DSL
modem/router in a single device: usually in this case you can set the
router to "bridge" mode somehow and plug your own router in.
-
If you have a cable modem, as far as I can determine by searching
the Internet, the previous does not apply to you. You can put a Virgin
Media Hub into modem mode,
but the effect is to turn off NAT and make the DHCP serve your
external IP address - so not the same, really.
So the tl;dr is that PPPoE is an important use case for a home
router - but if you have to order in a second fibre line to the
property to try it, that rather reduces the opportunities for people
to hack on it. So this is why I've been working on
bordervm - a QEMU VM
I can use with a passthru PCI ethernet card, that provides a PPPoE
service and relays it to an L2TP (Layer 2 Tunneling Protocol)
somewhere on the Internet. Mine is from Andrews & Arnold but presumably
others also exist.
Writing Things Down
The Liminix README file was getting very long and quite unstructured -
and it seemed unlikely that trend would reverse. So I've started on
writing a manual.
Tangential developments
I also upgraded my self-hosted Pleroma service from 2.4.4. to 2.5 and
it seems to be a lot faster and less timeout-y (here's hoping it stays
that way). This isn't Liminix-related per se, but sort of relevant
in that a lot of what I post about is Liminix.
@dan@brvt.telent.net
Order, Order#
Sun 26 Feb 2023 11:25:12 AM UTC
tl;dr for this week: I can browse the internet on an Android phone
connected by wifi to my Liminix test device :-)
Read on for how we got there.
Wireless
For the moment, wireless support is provided using the same pattern as
NixWRT was using (originally copied from OpenWrt):
-
we are using the same kernel version as OpenWrt, and the multitude
of patches they apply for wider device support. Usually this lags the
mainline kernel by a couple of minor versions, because of all those
patches.
-
this kernel is not configured for wifi
-
we build mac80211 and everything it depends on as modules from a
more recent kernel version, using the kernel Backports
Project which
automatically patch the more recent kernel sources into something that
will be compatible with the OpenWrt kernel.
It has to be said that this may be over-engineered for the current
state of the OpenWrt kernel. The latest kernel version that backports
can translate from is 5.15.92, and the OpenWrt kernel is 5.15.71, so
... not a whole lot of gain. Maybe the backports project will start
supporting 6.x kernels and make this worthwhile. I'm maintaining a
watching brief to see if we can remove this step, but the code was
mostly copied from NixWRT anyway so it didn't really take that much
longer than the simple approach would have done.
Anyway, having NixWRT code to copy/adapt here made it comparatively
quick to reach the point of "I can see the device with a wifi scanner"
Everything else
To get a device from "shows up in a wifi scan" to "can connect to and
reach the internet" needs a few more ingredients: it has to be able to
give connecting peers their IP addresses, provide them with name
service, forward packets, perform NAT for IPv4 etc.
rotuer.nix
is the configuration file I created for my test device to do all of
the above. As you may note from the comments at the top of the file,
it's still spike/experimental quality: less a preferred solution and
more an exploration of the problem space.
DHCP and DNS service for the LAN
We're using the excellent dnsmasq software which does both of these in
one binary. It's not a fully fledged recursive DNS server though, it
needs to know where to send requests upstream. Therefore we also need
to:
Make PPPoE service request DNS server addresses from upstream
We add the pppd option usepeerdns, then the ip-up script is called
with $DNS1 and $DNS2 environment variables. We write these into
the service outputs.
Resolvconf service
The days followed one another patiently. Right back at the
beginning of the multiverse they had tried all passing at the same
time, and it hadn’t worked. (Terry Pratchett, Wyrd Sisters)
Ordering constraints were a recurring theme of this week's work.
We need to have DHCP service running even when the WAN is not up.
Otherwise, imagine you have an internet outage and can't login to the
router to look at the logs and find out what's wrong because the
router won't give your device an IP address.
So, dnsmasq needs to start before PPPoE is running. But it also needs
to have the upstream nameserver addresses when PPPoE is running.
Happily, it does 90% of the required work to make this happen: if
configured with -resolv-file=/path/to/resolv.conf and if the
directory containing that file exists, it will watch the directory
(using polling or inotify or something) and read the file whenever
it's created or updated.
The missing piece is to create that file, then. We do this in a
service called resolvconf which depends on the pppoe service. It
needs to set file permissions that the dnsmasq service can read, so
this was a good time to introduce a bit of support/infra (a shell
functions file that service scripts can source) to make it easy for
services to get this right.
Bridge wired LAN with wireless
Not strictly necessary for wireless alone, but there are also LAN
ports on the test router which I'd like to be able to use. it turns
out you can't add a device to a bridge unless that device is both
"administratively up" and "operationally up", which is a poser for
wlan0 because it's not operationally up until hostapd has started up
and done various bits of twiddling. hostapd doesn't have readiness
notifications, so we can't rely on it to tell us when the interface is
ready. Did I say that ordering constraints were a theme this week?
What we do here is write a small program that listens for
Rtnetlink
events, and sends an s6 readiness notification when it sees that
wlan0 is ready to use. Then we can have a
service depending on it that adds the device to the bridge.
Packet forwarding
Is a simple matter of twiddling files in /proc/sys/net/ipv4/
NAT for the LAN devices
I've decided to go all-in on nftables here instead of supporting the
legacy (apparently, now) iptables. I could really use some good
nftables documentation, but it currently looks like the only way to
get that would be to write it myself.
Next?
I could spend a lot more time on cleaning this script up and
incorporating it into Liminix proper - and on supporting IPv6, which I
think is a hard requirement for a civilised society - and at some
point I will have to. But: it's all outside scope for phase 1, which
is "get the hardware devices that worked with NixWRT to work again
with Liminix". So I probably should put that on the back burner for
the moment, and turn instead to:
- make the ath10k radio (needed for 5GHz) work in this gl-ar750
- find out if there's a switch for the LAN ports (I assume there is)
and make it work if so
- make the gl-mt300a and mt300n work
The latter two are both building successfully, but do they boot? Don't
know yet.