292
S. Antonopoulos et al.
46.1 Introduction
Research and development of the PROGNOS post-processing system started as an
answer to various operational needs that the current system (UMOS; [10]) has difficulty addressing. For example, adapting and validating UMOS to very frequent
numerical model (NWP) updates has become tedious and time consuming. Furthermore, UMOS has a limited ability to evaluate various statistical modeling and
data pre-processing methods as well as to post-process new predictands. In addition, accounting for new observation datasets in the current operational system has
proven difficult. Therefore, UMOS is limiting our ability to accommodate evolving
forecasting program requirements such as gridded post-processing. The proposed
modular and flexible design of the PROGNOS system will address these challenges
while enhanced portability and transparency in the underlying code will help mitigate
some of the ongoing support requirements.
46.2 Architecture
46.2.1 Technologies
PROGNOS is built on open source technologies in Python, R and Bash languages.
The data pre-processing and statistical post-processing components of the system are
programmed in R, which makes use of the extensive statistical R libraries and support
from the online community. The data flow (I/O) managing modules mostly use R with
a few Python-based tools, but a gradual and complete move towards Python tools is
in progress. Daily experimental runs of the system are conducted with Maestro, an
ECCC sequencer [5], which is an application that submits sets of tasks (job scripts)
in a user-defined order. To enable portability, usability and flexibility of the system,
a centralized configuration approach is currently under development using YAML
(YAML Ain’t Markup Language), a human-readable, structured data serialization
syntax.
To ensure that the system is isolated, portable and reproducible, the following
library management technologies are used: the R Packrat package [8] and the Simple
Software Manager (SSM), an ECCC-developed packaging solution [4].
Furthermore, point based observation, predictor and forecast data are currently
stored and managed using SQLite, a serverless relational database management system, with plans to move to PostgreSQL hosted on a dedicated server. Adopting
SQL technologies is an improvement to the proprietary standards used in the current
system. It enhances I/O flexibility and versatility in both assimilating and selecting
predictors as well as in stratifying observations according to station metadata.
S. Antonopoulos et al.
46.1 Introduction
Research and development of the PROGNOS post-processing system started as an
answer to various operational needs that the current system (UMOS; [10]) has difficulty addressing. For example, adapting and validating UMOS to very frequent
numerical model (NWP) updates has become tedious and time consuming. Furthermore, UMOS has a limited ability to evaluate various statistical modeling and
data pre-processing methods as well as to post-process new predictands. In addition, accounting for new observation datasets in the current operational system has
proven difficult. Therefore, UMOS is limiting our ability to accommodate evolving
forecasting program requirements such as gridded post-processing. The proposed
modular and flexible design of the PROGNOS system will address these challenges
while enhanced portability and transparency in the underlying code will help mitigate
some of the ongoing support requirements.
46.2 Architecture
46.2.1 Technologies
PROGNOS is built on open source technologies in Python, R and Bash languages.
The data pre-processing and statistical post-processing components of the system are
programmed in R, which makes use of the extensive statistical R libraries and support
from the online community. The data flow (I/O) managing modules mostly use R with
a few Python-based tools, but a gradual and complete move towards Python tools is
in progress. Daily experimental runs of the system are conducted with Maestro, an
ECCC sequencer [5], which is an application that submits sets of tasks (job scripts)
in a user-defined order. To enable portability, usability and flexibility of the system,
a centralized configuration approach is currently under development using YAML
(YAML Ain’t Markup Language), a human-readable, structured data serialization
syntax.
To ensure that the system is isolated, portable and reproducible, the following
library management technologies are used: the R Packrat package [8] and the Simple
Software Manager (SSM), an ECCC-developed packaging solution [4].
Furthermore, point based observation, predictor and forecast data are currently
stored and managed using SQLite, a serverless relational database management system, with plans to move to PostgreSQL hosted on a dedicated server. Adopting
SQL technologies is an improvement to the proprietary standards used in the current
system. It enhances I/O flexibility and versatility in both assimilating and selecting
predictors as well as in stratifying observations according to station metadata.
