I’m suffering from quite frequent (sometimes up to four a day — the N810 is nearly always on and for >90% of the time connected to the Internet (WLAN or UMTS-over-BT) reboots of my N810. Reading ReportingRebootIssues, I went to check the common files:
~ $ cat /proc/bootreason
32wd_to
~ $ ls -la /var/lib/dsme/stats/lifeguard_resets
ls: /var/lib/dsme/stats/lifeguard_resets: No such file or directory
~ $ ls -la /var/lib/dsme/stats/
drwxr-xr-x 2 root root 0 Jan 21 22:11 .
drwxr-xr-x 3 root root 0 Jan 1 1970 ..
-rw-r--r-- 1 root root 3 Jan 21 22:11 32wd_to
-rw-r--r-- 1 root root 61 Jan 17 04:32 lifeguard_restarts
-rw-r--r-- 1 root root 2 Jan 17 17:42 sw_rst
~ $ more /var/lib/dsme/stats/32wd_to
14
~ $ more /var/lib/dsme/stats/lifeguard_restarts
/usr/bin/hildon-input-method : 4
/usr/sbin/ke-recv : 4 *
Digging deeper down, I even checked dmesg (because of “If the device boots and reboots several times, then log in to a shell and check “dmesg” for whether this boot also ran out of memory (“OOM”) and the kernel happened to kill something that was not essential this time.”) — the only troublesome (for my liking) entry I found is:
[ 62.671875] JFFS2 notice: (403) check_node_data: wrong data CRC in data node at 0x01b87580: read 0xf56b695c, calculated 0xd4218b92.
Since I don’t use any MMC/SC yet, this must relate to the internal 2 GB of flash memory. I’m not that much into flash file-systems, does that mean I have a) broken internal flash memory, or b) there was data corruption (maybe due to the reboot(s)) or c) there is still a bug lurking around and I did trigger it? I’ve found a patch while googling around, which seems to be destined for 2.6.21 and tackle CRC issues:
commit 10731f83009e2556f98ffa5c7c2cbffe66dacfb3
Author: Artem Bityutskiy <Artem.Bityutskiy@nokia.com>
Date: Wed Apr 4 13:59:11 2007 +0300
[JFFS2] fix buffer sise calculations in jffs2_get_inode_nodes()
In read inode we have an optimization which prevents one min. I/O unit (e.g. NAND page) to be read more then once.
Namely, at the beginning we do not know which node type we read, so we read so we assume we read the directory entry, because it has the smallest node header. When we read it, we read up to the next min. I/O unit, just because if later we’ll need to read more, we already have this data.
If it turns out to be that the node is not directory entry, and we need more data, and we did not read it because it sits in the next min. I/O unit, we read the whole next (or several next) min. I/O unit(s). And if it happens to be that we read a data node, and we’ve read part of its data, we calculate partial CRC. So if later we need to check data CRC, we’ll only read the rest of the data from further min. I/O units and continue CRC checking.
This code was a bit messy and buggy. The bug was that it assumed relatively large min. I/O unit, so that the largest node header could overlap only one min. I/O unit boundary.
This parch clean-ups the code a bit and fixes this bug.
The patch was not tested on flash with small min. I/O unit, like NOR-ECC, nut it was tested on NAND with 512 bytes NAND page, so it at least does not break NAND. It was also tested with mtdram so it should not break NOR.
Signed-off-by: Artem Bityutskiy <Artem.Bityutskiy@nokia.com>
Signed-off-by: David Woodhouse <dwmw2@infradead.org>
My N810 tells me it’s running “2.6.21-omap1 #2 Fri Dec 7 11:17:13 EET 2007 armv6l unknown” – so, am I safe?