SatHDSS
Network panel

Algae Dataset

Every data hub, dataset, sensor, variable, paper and limitation in one tree — and for each: where it comes from, how far back it goes, how often it captures, why it belongs in a cholera study, and why it does not work everywhere.

Highlight site All three Chakaria HDSS Matlab HDSS Dhaka City Choosing a site greys out everything that does not apply there.

Overview

Algae Dataset
Chakaria (coastal) · Matlab (riverine) · Dhaka (urban inland) · 2000-01-01 → 2026-06-30
26
datasets
25
free
14
no account
13
sensors
7
data hubs
33
limitations
Reliable · accessible · authentic. Every dataset in this tree is free and comes from an operational provider — ESA, NASA, USGS or Copernicus. Authenticity is not the differentiator. What differs is resolution, revisit interval, and whether the water body at that site is physically large enough to be seen at all.
The limitation that governs everything. A 4 km ocean-colour pixel is 16 km². The widest reach of the Meghna is 6 km and every Dhaka lake is under 1 km across. That single fact is why sea-water products serve Chakaria directly, serve Matlab only as regional forcing, and do nothing at all for Dhaka.

Study-level limitations (6)

CRITICAL No in-situ chlorophyll anywhere in the project validation

Without paired field measurements, NDCI, FAI and the phycocyanin proxy are ordinal indices, not concentrations. Reviewers will treat every inland number as uncalibrated. This is the single largest weakness in the study and it is fixable comparatively cheaply.

What to do: Three asks, cheapest first: BORI Cox's Bazar coastal monitoring, Department of Fisheries shrimp-pond water-quality logs, and the open GLORIA hyperspectral archive. Even 30-50 paired observations move three products from tier C to tier A.

CRITICAL Outcome data is behind icddr,b ethical approval access

Cholera case series for all three sites require Research Review Committee and Ethical Review Committee approval plus a data-sharing agreement. Free for collaborators, but it takes weeks to months. Every satellite task can run in parallel · none of them can finish without it.

What to do: Submit the application before starting anything else. Ask for three things in one email: the Matlab village-ID (vid) lookup, the Chakaria union-level shapefile, and whether Chakaria records lab-confirmed cholera or only diarrhoeal disease.

HIGH Monsoon cloud destroys June-September optical coverage atmospheric

Every optical sensor here - Sentinel-2, Sentinel-3, Landsat, MODIS, ocean colour - is blinded by cloud. June to September is the wettest period in Bangladesh and also the period of highest cholera transmission, so the data is thinnest exactly when it matters most. Worse, cloudiness is not random with respect to blooms: cloudy weeks have different bloom dynamics, which makes the missingness informative rather than ignorable.

What to do: Carry Sentinel-1 SAR water extent as a cloud-proof companion, and model cloud_frac explicitly. The hospital codebook already contains ground sunshine hours and cloud cover (rows 306-307) - use them to characterise what you lost.

HIGH The 2015 sensor discontinuity temporal

Sentinel-2 begins mid-2015 and Sentinel-3 late 2016. Before that you have Landsat at 16 days and MODIS at 250 m. Blending sensors without correction creates a step change in 2015 that a model will read as a trend - a completely artefactual one.

What to do: 11_algae_daily_fill.py regresses every sensor onto a reference on same-day overlaps and writes the coefficients to algae_daily_bias_coeffs.csv. Any pair with fewer than 20 overlaps is left UNCORRECTED and flagged rather than silently fudged.

HIGH Suspended sediment biases every inland chlorophyll retrieval algorithmic

The Meghna carries an enormous sediment load and Dhaka's water bodies are turbid year-round. High TSS inflates red and NIR reflectance, which contaminates NDCI and FAI. This is measurement error, not collinearity - the chlorophyll series is partly a sediment series.

What to do: Always extract turbidity alongside chlorophyll and carry it as a covariate whether or not it is a predictor of interest. It is in the pipeline output as the turbidity column.

MEDIUM Multiple testing across the lag grid coverage

34 variables x 14 lag steps x 3 sites is 1,428 cross-correlations. At p<0.05 roughly 70 will clear significance by chance alone, and the strongest of those will look publishable.

What to do: Pre-register the expected lag per variable before looking at the heatmap. Treat the rest of the surface as exploratory and say so.