HR: 16:30h
AN: H54B-03 INVITED [Abstracts]
TI: Neural Network Hydrological Modelling: Linear Output Activation Functions?
AU: * Abrahart, R J
EM: bob.abrahart@nottingham.ac.uk
AF: School of Geography, University of Nottingham, Nottingham, NG7 2RD
United Kingdom
AU: Dawson, C W
EM: c.w.dawson1@lboro.ac.uk
AF: Department of Computer Science, Loughborough University, Loughborough, LE11 3TU
United Kingdom
AB:
The power to represent non-linear hydrological processes is of paramount importance in neural network hydrological modelling
operations. The accepted wisdom requires non-polynomial activation functions to be incorporated in the hidden units such that
a single tier of hidden units can thereafter be used to provide a 'universal approximation' to whatever particular
hydrological mechanism or function is of interest to the modeller. The user can select from a set of default activation
functions, or in certain software packages, is able to define their own function - the most popular options being logistic,
sigmoid and hyperbolic tangent. If a unit does not transform its inputs it is said to possess a 'linear activation function'
and a combination of linear activation functions will produce a linear solution; whereas the use of non-linear activation
functions will produce non-linear solutions in which the principle of superposition does not hold. For hidden units, speed of
learning and network complexities are important issues. For the output units, it is desirable to select an activation
function that is suited to the distribution of the target values: e.g. binary targets (logistic); categorical targets
(softmax); continuous-valued targets with a bounded range (logistic / tanh); positive target values with no known upper bound
(exponential; but beware of overflow); continuous-valued targets with no known bounds (linear). It is also standard practice
in most hydrological applications to use the default software settings and to insert a set of identical non-linear
activation functions in the hidden layer and output layer processing units. Mixed combinations have nevertheless been
reported in several hydrological modelling papers and the full ramifications of such activities requires further
investigation and assessment i.e. non-linear activation functions in the hidden units connected to linear or clipped-linear
activation functions in the output unit. There are two obvious advantages related to the use of a linear activation function
in the output unit: (i) to restrict potential impacts and distortions associated with upper limit and lower limit saturation
effects; and (ii) to address potential deficiencies and ceilings associated with undershoots or requirements to extrapolate
beyond the range of the training dataset. The harmful side effects of using linear as opposed to non-linear activation
functions in the output unit will be reported in this paper based on an investigation of six-hour timestep operational river
level forecasting for the Skelton Gauging Station [Station No: 027009; Grid Ref: SE 568 554] on the River Ouse in England.
The power to develop simple near-linear one-step-ahead forecasts remained more or less unchanged; whereas the challenge to
develop more demanding non-linear four-step-ahead forecasts revealed major shortcomings related to the implementation of a
linear activation function in the output unit of a parsimonious neural network model.
DE: 1805 Computational hydrology
SC: Hydrology [H]
MN: Fall Meeting 2005