46 PROGNOS: A Meteorological Service of Canada (MSC) Initiative …
293
Fig. 46.1 Representation of the PROGNOS flow chart (future developments in green)
46.2.2 Infrastructure and Flow Chart
As shown below (Fig. 46.1), the proposed system consists of three main components:
(1) Data ingestion modules that fill the PROGNOS database with observations as well
as calculated and/or interpolated predictors; (2) the PROGNOS core consisting of
various modules responsible in part for the pre-processing of the input data, the
training of the forecast models and the production of diagnostics and forecasts; and
(3) the output module, which holds the forecast database and is also responsible for
converting the forecasts to a format compatible with SCRIBE, the current public
forecast dissemination system [9].
46.3 System Features
Several post-processing approaches such as LASSO [7], random forest [1] and
Kalman filter [3] have been explored in previous versions of the system; however,
most of the recent development has been centered on multiple linear regression
(MLR) techniques. A batch calibration approach is currently implemented and executed weekly to update model coefficients. Moreover, capabilities for updatable
model output statistics are being considered. The predictor selection scheme is a
stepwise approach, which consists of sequential replacement [2] using a weighted
Bayesian information criterion (BIC; [6]). The BIC allows for the comparison
between candidate MLR equations in order to select the best fit model for each
observation station and forecast hour. The possibility of adding weighted transition
schemes to adjust for seasonal changes and numerical model updates is accounted
for.
The training, data pre-processing and the forecasting modules have the capability
of running in parallel to make use of the full potential computational resources
available. The experimental runs are using a simple multi-node approach to distribute
the processing of each data block to different CPUs. Additional code optimisation
293
Fig. 46.1 Representation of the PROGNOS flow chart (future developments in green)
46.2.2 Infrastructure and Flow Chart
As shown below (Fig. 46.1), the proposed system consists of three main components:
(1) Data ingestion modules that fill the PROGNOS database with observations as well
as calculated and/or interpolated predictors; (2) the PROGNOS core consisting of
various modules responsible in part for the pre-processing of the input data, the
training of the forecast models and the production of diagnostics and forecasts; and
(3) the output module, which holds the forecast database and is also responsible for
converting the forecasts to a format compatible with SCRIBE, the current public
forecast dissemination system [9].
46.3 System Features
Several post-processing approaches such as LASSO [7], random forest [1] and
Kalman filter [3] have been explored in previous versions of the system; however,
most of the recent development has been centered on multiple linear regression
(MLR) techniques. A batch calibration approach is currently implemented and executed weekly to update model coefficients. Moreover, capabilities for updatable
model output statistics are being considered. The predictor selection scheme is a
stepwise approach, which consists of sequential replacement [2] using a weighted
Bayesian information criterion (BIC; [6]). The BIC allows for the comparison
between candidate MLR equations in order to select the best fit model for each
observation station and forecast hour. The possibility of adding weighted transition
schemes to adjust for seasonal changes and numerical model updates is accounted
for.
The training, data pre-processing and the forecasting modules have the capability
of running in parallel to make use of the full potential computational resources
available. The experimental runs are using a simple multi-node approach to distribute
the processing of each data block to different CPUs. Additional code optimisation
