HR: 16:30h
AN: H54B-03 INVITED     [Abstracts]
TI: Neural Network Hydrological Modelling: Linear Output Activation Functions?
AU: * Abrahart, R J
EM: bob.abrahart@nottingham.ac.uk
AF: School of Geography, University of Nottingham, Nottingham, NG7 2RD United Kingdom
AU: Dawson, C W
EM: c.w.dawson1@lboro.ac.uk
AF: Department of Computer Science, Loughborough University, Loughborough, LE11 3TU United Kingdom
AB: The power to represent non-linear hydrological processes is of paramount importance in neural network hydrological modelling operations. The accepted wisdom requires non-polynomial activation functions to be incorporated in the hidden units such that a single tier of hidden units can thereafter be used to provide a 'universal approximation' to whatever particular hydrological mechanism or function is of interest to the modeller. The user can select from a set of default activation functions, or in certain software packages, is able to define their own function - the most popular options being logistic, sigmoid and hyperbolic tangent. If a unit does not transform its inputs it is said to possess a 'linear activation function' and a combination of linear activation functions will produce a linear solution; whereas the use of non-linear activation functions will produce non-linear solutions in which the principle of superposition does not hold. For hidden units, speed of learning and network complexities are important issues. For the output units, it is desirable to select an activation function that is suited to the distribution of the target values: e.g. binary targets (logistic); categorical targets (softmax); continuous-valued targets with a bounded range (logistic / tanh); positive target values with no known upper bound (exponential; but beware of overflow); continuous-valued targets with no known bounds (linear). It is also standard practice in most hydrological applications to use the default software settings and to insert a set of identical non-linear activation functions in the hidden layer and output layer processing units. Mixed combinations have nevertheless been reported in several hydrological modelling papers and the full ramifications of such activities requires further investigation and assessment i.e. non-linear activation functions in the hidden units connected to linear or clipped-linear activation functions in the output unit. There are two obvious advantages related to the use of a linear activation function in the output unit: (i) to restrict potential impacts and distortions associated with upper limit and lower limit saturation effects; and (ii) to address potential deficiencies and ceilings associated with undershoots or requirements to extrapolate beyond the range of the training dataset. The harmful side effects of using linear as opposed to non-linear activation functions in the output unit will be reported in this paper based on an investigation of six-hour timestep operational river level forecasting for the Skelton Gauging Station [Station No: 027009; Grid Ref: SE 568 554] on the River Ouse in England. The power to develop simple near-linear one-step-ahead forecasts remained more or less unchanged; whereas the challenge to develop more demanding non-linear four-step-ahead forecasts revealed major shortcomings related to the implementation of a linear activation function in the output unit of a parsimonious neural network model.
DE: 1805 Computational hydrology
SC: Hydrology [H]
MN: Fall Meeting 2005