Wednesday, August 15, 2012

info from brief resumptions of communications

In the last few days, from August 10th - 13th, there have been four brief periods when communications with the Saipan station have resumed.  This post explains what we have learned from the existence and the contents of those communications.

First let's look at the timing and duration of these periods, speaking in local Saipan time:
  • Friday, August 10th, 7:18 am - 7:42 am (5 records)
  • Saturday, August 11th, 1:30 pm - 4:48 pm (34 records)
  • Sunday, August 12th, 11:24 am - 11:30 am (2 records)
  • Monday, August 13th, 9:42 am - 1:54 pm (33 records)
 The record counts I've listed are from the 6-minute data table.

First of all, note that every record in a data table is given a sequential record number.  This makes it possible to identify cases where records are (or are not) consecutive.  Because of this we can say with high probability that the station's datalogger was completely non-functional from the period from July 19th to August 9th (with times in UTC).  There is no chance that this could be yet another communications-only failure, with little impact otherwise to station operations.

The pattern of communications is somewhat consistent with the explanation that, following a severe power loss that to some extent drained the station's rechargeable batteries, the station may be recharging itself and may briefly resume operations during daylight hours.  However there is still no explanation for the initial power loss, or for the previous power drops (May 19th - June 4th, June 26th - July 5th) seen in the data record.

Perhaps worse, this pattern suggests that the intermittent voltage drop may not be over.  Normally if a short-circuit is repaired we would expect to see better and longer "ontimes" with each passing day.  Instead we start with a brief uptime in the early morning (followed by silence during that day's prime daylight hours).  We also see a stronger performance Saturday followed by a weaker performance Sunday.  Monday's communications are longer but they describe power levels that are getting lower throughout the morning and early afternoon.  And on Tuesday, which has already ended in Saipan, there was no resumption of communications at all.

Another problem is that many of the station's instruments are now malfunctioning.  This may be due to electrical problems (an artifact of running them at very low levels) or their calibration files and settings may have become corrupted, leading them to produce reports in a format that the datalogger program was not designed to parse.  A brief rundown of the state of station instruments follows:

The standalone air temperature sensor, barometer, anemometer and electronic compass all appear to be working normally.  These are all analog instruments and do not depend on serial communications in any way.

Both light sensors are damaged.  The surface light sensor continues to produce serial reports of some kind (as evidenced by that sensor's instrument "counts" in the logger) but these reports are apparently filled only with zeroes, day or night, including for temperature and voltage.  The underwater light sensor has been offline since last October and this has not changed.

The Deep CTD (Teledyne) seems to be at least partially operational.  Its conductivity and temperature readings seem reasonable, although its depth reading show an odd 30-cm shift in the past few days and may indicate a problem.

The PacIOOS CTD was producing data reports on Friday and Saturday but since that time its output has been in a format that was unrecognized by the datalogger programming.  It produced a full "status" report (in response to hourly prompting by the datalogger) on August 11th, 2012, at 4:02 UTC.  This report's contents were as follows:
  • Year: 2012
  • Month: 8
  • Day: 11
  • Hour: 4
  • Minute: 2
  • Second: 56
  • Serial Num: 1606481
  • Num Events: 15
  • Volts Main: 7.3
  • Volts Lith: 8
  • Curr Main: 61.3
  • Curr Pump: 283.4
  • Curr Ext: 286.8
  • Mem Bytes: 1082107
  • Samples: 56953
  • Sample Free: 3406107
  • Sample Len: 19
  • Headers: 4
The Vaisala WXT, like the surface light sensor, is apparently producing reports but as recorded by the datalogger these reports are all zeroes.

It is possible that there has been damage to the datalogger, memory unit, serial port units, radio or cellular modem, although so far there is no sign of such damage.

The main problem at this point is not the instrument failures but that of isolating the cause of these voltage drops.  It seems like there have been voltage drops since mid-May, that a severe voltage drop has rendered the station non-functional for three weeks, and that this may be an intermittent problem that is still ongoing.  If we knew or suspected the cause of the power issues then it would simply be a matter of replacing the station instruments with the spares in storage at CRM.  However our first focus must be on diagnosing the underlying cause of the power losses and then repairing it.

Mike J+

Tuesday, August 7, 2012

Maintenance log: Launched CRM vessel from shore

[From David Benavente's LAOLAO Bay ICON Station Maintenance log -- August 6, 2012:]

Launched CRM vessel from shore. Station was climbed by David B. Upon opening the brain housing all wires and connection plugs were first inspected. No sign of damage was found. Next, Brain hardware was visually inspected. Data logger lights, modem lights, and power supply lights were all illuminated. The inspection did not reveal any wiring problems to the brain.
Underwater cleaning and maintenance of station instruments was also conducted, during this visit.


Tuesday, July 24, 2012

Station Power Failure

The Saipan CREWS station went offline as of 22:36:35 UTC on Thursday, July 19th, 2012.  In local Saipan time this is Friday morning, July 20th, at 8:36 am.  Miami time, it is Thursday night, July 19th, at 6:36 pm.

Gordon Walker (who is taking Ross Timmerman's position with PacIOOS) sent out this information on Monday, July 23rd:
We were able to connect to the modem at 9:30 am (locate Saipan time) and everything looked okay except for the voltage.  It is at 5.76 V.  It seems as though the power for the entire array is extremely low.  This may be the reason for the interruption.  We will continue to monitor the modem on our end.
Unlike previous outages, which were believed to be primarily caused by modem or cellular network problems, this latest outage is clearly due to a loss of power on the station.  Consider the graph of minimum (blue), average (red) and maximum (green) hourly datalogger voltages, plotted against Julian Day, for all of 2012, at right (click on the image for a larger version).

A CREWS station's normal operating voltage cycles between 12.5 V and 14 V, with peak voltages during the sunlight hours when the solar panels are supplying power and lowest voltages overnight when they are not.

A close examination of this station's power levels in 2012 shows three periods where the station did not appear to be reaching its usual daylight voltage peaks.  The first occurred from roughly May 19th to June 4th; the second from June 26th to July 5th; and the third began on July 12th and ended with the station's loss of power.

That third drop, where the station power level fell below 10 V before communications with the station entirely ceased, has happened before in CREWS history, though it is rare.  It happened once in Jamaica, in June/July 2008, when that station's "deep" light sensor's bulkhead connector was broken (we think this was caused by a period of strong currents and the sensor's cable being insufficiently tied down).  In that case the failed sensor was removed after two days, and the station resumed communications about 14 days later when its batteries were again sufficiently charged.  More or less the same thing happened in Puerto Rico, twice, once in April of 2010 and once in July of 2011.  Both of these latter outages were caused by flooded light sensors, at least one of which was caused by a puncture in the sensing surface of the instrument.

Before I continue, I'll just emphasize that what happened in Saipan on July 19th is nothing like the outages that this station has experienced before -- the first failure on October 2nd, or the more recent outages (April 12th - 19th, those few days in early May, and May 12th - 29th).  Those were quite clearly caused by problems with the cellular modem or the Docomo network.  They were communications outages, only.  After communications resumed, the data record showed that the station continued to operate (and store data locally) after we lost contact.  [In the October incident, the station continued to operate only for two more days, when a blown fuse up top took it out completely.]  In this case, to be clear, we believe all station functions to be non-operational, with the possible exception of the battery-powered PacIOOS CTD.

However, this current outage is somewhat different from the power losses we've seen before at Jamaica and Puerto Rico, for two reasons.  In the current case, there is a strange history of lower voltage levels for days or weeks at a time, which were then followed (the first two times) by a return to normal power levels.  This current problem, it seems, is somehow intermittent, which does not lend itself to explanation by something irreversible like a flooded instrument.

The second difference here is that those previous incidents left enough clues for us to determine with high probability what had failed.  Specifically, in the previous cases there was a measured voltage drop in one instrument only, which was then followed (in a matter of days, or hours) by the station's complete power loss.  This current Saipan power loss does not include any such hints in the power levels of its instruments.  This may be because one of those instrument (the underwater light sensor, as it happens) is already incommunicado.  Perhaps that light sensor has a loose wire in its communications but continues to take a full power/ground feed from the station.  If so, it could flood without warning and cause the failure of the entire station as happened in Jamaica and Puerto Rico.

However, those previous periods of voltage loss and recovery remain unexplained, and this suggests that a flooded instrument may not turn out to be the cause in this case.

Mike J+

Friday, July 6, 2012

Maintenance log: clean SeaBird lens cleaner

[From David Benavente's LAOLAO Bay ICON Station Maintenance log -- July 5, 2012:]

Objective was to revisit the station and thoroughly clean SeaBird lens cleaner. David B had a scheduling conflict he joined the group later. On this visit Copper screens were replaced. They appeared to be punctured; the puncture marks looked as if they had come from a three prong spear, commonly used by fishermen.

Wednesday, June 27, 2012

Maintenance log: Accessed Station through cliff

[From David Benavente's LAOLAO Bay ICON Station Maintenance log -- June 26, 2012:]

Accessed Station through cliff. David B. used SCUBA, while Rod C., John I., and Steven J. snorkeled. Found some evidence that fishermen were in close vicinity to the station. Also a knife was found below station, this could be because the base of the station appears to have become a congregation area for Trochus mussels. Because these are a harvested species fishermen may be visiting the station more often.

Saturday, June 23, 2012

Docomo cellular connection spotty, error-prone

Following the recent modem outage (May 12th - 29th), I attempted my usual practice of a gradual re-enabling of the data table downloads.

A note about data tables:  the datalogger has six data tables, each with its own collection schedule.  One very simple table stores one record per day, with only a few data values that are used to monitor the memory card's diagnostics and confirm that it is formatted/storing correctly.  At the other end of the spectrum, one data table stores a new record every five seconds, to record each individual reading from the analog sensors (air temperature and barometric pressure).

Historically most of the CREWS data record has come from the one-hour data table, since most CREWS stations have relied on once-hourly, twenty-second windows of communication on the GOES East satellite.  Data were summarized as necessary and transmitted in a format designed to fit comfortably in our communications window.

With the advent of larger-capacity memory cards, we began to store more granular records in the datalogger's memory for later retrieval.  With our "always on" connection to the Saipan CREWS station (apart from the cellular network problems, that is), we began to access all of those data in real time.

As it might be expected, the 5-second data table contains a very large number of records, although each record stores relatively few data values.  There is also a 30-second, a 1-minute, a 6-minute, a 1-hour and (as mentioned before) a 1-day data table.

Since our data feeds require only the 1-hour and 6-minute data tables for processing, I generally disable the download of all other data tables during long modem outages, in order that when service resumes our feeds can come up as quickly as possible.  Later I will manually re-enable download of the larger data tables and monitor their progress to ensure that our main data feeds aren't negatively impacted.  This was more or less how things worked following the April 12th - 19th outage earlier this year, with all data tables recovered by April 24th.

However, after the May 29th re-establishment of communications with the modem, two differences were observed.  For one, the LoggerNet monitoring software was reporting a large number of error conditions and failed connections, although it did not seem to be impacting the download of the 1-hour or 6-minute data tables.  What was worse, however, is that attempts to re-enable the download of the other data tables appeared to be knocking the modem entirely offline, requiring a request to Docomo for modem/network reset.

The first happened on the morning of Wednesday, May 30th, 2012.  I re-enabled the first of the data tables for download, and all communications ceased.  I alerted Ross and he replied later that same day:
I sent another email to Luigi [at Docomo] this morning to reset the modem. I will take down the modem settings and have Sierra Wireless take a look. It isn't clear whether the Docomo network or the device is the cause of these outages. No reports of network issues have been received, but as David mentioned, they can be frequent. I don't quite understand how we had good service for the initial months and not now.
This apparently got a response from Docomo immediately and the modem was back online again the next morning.  However, it once again went offline when I attempted to re-enable download of the other data tables.  From my message to Ross on May 31st:
You might want to check the modem again today!  It seems like it might be locking up whenever I try to pull data files from the logger.  I don't know why it would suddenly be reacting this way because this was S.O.P. last year (August/Sept) and during the uptimes earlier this year (March 19th through May 12th). And most notably on April 24th, when I did some very intensive data transfers from the modem and everything seemed fine.
This message may have caught Ross away from his email, for his reply arrived on June 4th (wherein he said he would contact Docomo) with an update on June 5th (to report that the modem was online again).

Following this clear pattern of outages, I did not attempt to access the other data tables right away.  Instead, I allowed the main 1-hour and 6-minute data tables to run normally and populate our data feeds.  However, several weeks later I glanced at the LoggerNet status readouts and noticed that the pattern of error conditions and failed connection for the Saipan modem appeared to have resolved itself.  All communications appeared to be normal (for example, in comparison with our other cellular modem which is deployed at Port Everglades near Fort Lauderdale, FL).

When I noticed this change, I re-enabled download of the other data tables and this time, they all downloaded beautifully.  This caught us up back to the beginning of the May 12th outage for the first time, and we once again were keeping current with all data tables in real time.  I sent out this by email on Friday, June 22nd:
LoggerNet's error rate on the docomo modem suddenly relaxed this past week and I took a chance and re-enabled the download of those more granular files I'd mentioned.  That worked beautifully, so we're all caught up with the missing data back to May 12th, and all files are once again downloading in realtime. I don't know what was causing the modem to cycle offline before but it seems to be okay for now.
Mike J+

Friday, June 1, 2012

Maintenance log: May maintenance

[From David Benavente's LAOLAO Bay ICON Station Maintenance log -- May, 2012:]

I was away on travel for most of May. To my understanding the MMT performed regular maintenance on the ICON station during this time. Steven J. reported that the station and its instruments were cleaned.