It has been shown by a number of people that the weights of such an
element are assured of arriving at a set of values which produce the
correct output-input relation after a sufficient number of errors,
provided that such a set exists. An upper bound on the number of
possible errors can be given which depends only on the initial weight
values and the logical function to be learned. This does not, however,
solve our network problem for two reasons.
First, as the number, n, of inputs gets large, the number of errors
to be expected for most functions which can be learned increases to
unreasonable values. For example, for n = 6, most such functions
result in 500 to 1000 errors compared to an average of 32 errors to be
expected in a perfect learning device.
Second, and more important, the fraction of those logical functions
which can be generated in a single element becomes vanishingly small as
n increases. For example, at n = 6 less than one in each three trillion
logical functions is so obtainable.
NETWORKS OF ELEMENTS
It can be demonstrated that if a sufficiently large number of linear
threshold elements is used, with the outputs of some being the inputs
of others, then a final output can be produced which is any desired
logical function of the inputs. The difficulty in such a network lies
in the fact that we are no longer provided with a knowledge of the
correct output for each element, but only for the final output. If the
final output is incorrect there is no obvious way to determine which
sets of weights should be altered.
As a result of considerable study and experimentation at Aeronutronic,
a network model has been evolved which, it is felt, will get around
these difficulties. It consists of four basic features which will now
be described.
Positive Interconnecting Weights
It is proposed that all weights in elements attached to inputs which
come from other elements in the network be restricted to positive
values. (Weights attached to the original inputs to the network, of
course, must be allowed to be of either sign.) The reason for such a
restriction is this. If element 1 is an input to element 2 with weight
c₁₂, element 2 to element 3 with weight c₂₃, _etc._, then the sign of
the product, c₁₂c₂₃ ..., gives the sense of the effect of a change in
the output of element 1 on the final element in the chain (assuming
this is the only such chain between the two elements). If these various
weights were of either possible sign, then a decision as to whether or
not to change the output in element 1 to help correct an error in the
final element would involve all weights in the chain. Moreover, since
there would in general be a multiplicity of such chains, the decision
is rendered impossibly difficult.
The above restriction removes this difficulty. If the output of any
element in the network is changed, say, from -1 to +1, the effect on
the final element, if it is affected at all, is in the same direction.
Public-domain text, read in full here on John Shaqi.
Reviews
Reviews
No reviews yet
Be the first to share your thoughts on this work.
Join the Discussion
Join the discussion
Sign in to leave a comment or review.
Sign InorCreate an account