Hello
I spent several hours yesterday troubleshooting a problem on my own home network. In the end I found the solution but it’s the root cause that I find very interesting.
I have a very simple network at home: Play ISP modem (fiber), then a Play router, my cheap Mercusys 8 port switch (because it’s the only switch small enough to fit into my garage telecommunications cupboard, and finally a mesh of 3 Mercusys H70x mesh access points. 2 APs are on the ground floor, 1 AP is on the first floor.
It all worked fairly well until a few days ago, when I started noticing some packet loss and very low throughput. I thought this was caused by the fact that I moved things around in my room and the distance between my PC and the AP on the first floor is a corridor + room walls + additional wall with a cupboard.
Now, there’s a wall socket in my room but it never worked because the switch in the garage only has 8 ports, and there are more sockets in my house. So in step 1 I dug up an old Cisco 891 router with 8 switch ports, i configured them all in the same vlan, set up an SVI of 192.168.1.240 on vlan 10, cleaned up some old lab config, ran a cable between the switch and the 891, connected the four cables that never fit into the small switch into the 891 router and the socket in my room on the first floor was now functional. But the throughput was still low/fluctuating between 4mbit/s and 400mbit/s. Not great.
In step 2 I called my ISP hotline and reported a problem with packet loss and low throughput. The lady asked my to run some tests and she was even clever enough to ask me to disconnect from wifi and connect directly via cable. I proposed an extra step of bypassing the switch and connecting straight to the Play router. Still no change, i got 8mbit/s. So i said to them that i’m going to restart their modem and router. The lady said fine, but in her opinion this wasn’t going to change anything. I found this strange because restarting everything is the best solution in any engineer’s handbook. So i did that, and the tests showed 500mbit/s immediately after the router came back to life. Wow. Problem solved, right?
Not so fast. Here’s the fun part. Now my Mercusys mesh was blinking all red. I restarted all 3 APs but then all my home endpoints lost internet access. I was very confused now. I tried to ssh to the 891 router and found that i couldn’t. Did i forget about anything?
I went to the garage, consoled into the router, found a missing line in the admin ACL on the vty line (allowing admin access from the home network 192.168.1.0/24), pinged 8.8.8.8 from the SVI on the home vlan, all good. I went back to my PC and there was still no internet access. What’s going on here. .. and why does my PC have a 192.168.68.25 address? where does this come from??? The play router should give out addresses in the 192.168.1.0/24 space. Who gives out those other addresses?
I went back to the garage, checked the ios config (an old DHCP pool maybe?), but there was nothing there. So i checked ARP cache on the PC and used the mac address of the 192.168.68.1 address in the mac vendor database. It’s Mercusys! but why?
Right, I guess now it’s time to switch of all 3 APs to see what happens. I switched them off, did ipconfig /renew on a PC… and got absolutely nothing. Oh. This meant that the DHCP server on the play router was dead. I restarted their router again, turned on APs… and got 192.168.68.24 again. Darn. I set a manual IP address, 192.168.1.201 and I got my internet connection back.
Now, because it was very late already i was too tired to think about WHY the mercusys APs started to give out addresses. Or maybe I didn’t think about this because another problem manifested itself: the AP mesh was flapping all the time. Ideally, the first AP should be selected as MAIN and the other ones should be “wired connected” via the wire backhaul (through the garage switch) to each other. But the status was changing all the time from “wired” to “wifi mesh” and finally the other non-main APs would go offline.
After wasting another hour reading through reddit forums (where users reported similar problems with the flapping mesh), I decided to go to Mercusys documentation site. What caught my attention was the fact that they can operate in router mode and in this mode they cannot be connected to each other via a switch, they can only be daisy chained, OR they can operate in the mesh network mode, with a wire backhaul through a switch. This was perhaps some clue. In the Mercusys application i had to go into a quite well hidden option that changed the mode from Router to Mesh and it all went green after a few minutes.
So what happened here? Apparently, as soon as the Play router’s DHCP failed, Mercusys APs couldn’t connect to the cloud anymore and switched automatically to the router mode and started to give out addresses in their own address space even though the device should know that all packets will be blackholed because they couldn’t get to the internet anyway. This flawed “cascading” effect of the ISP DHCP failure cost me almost 5 hours of troubleshooting even though I have an extremely simple home network.