The Perceptron: A Probabilistic Model for Information Storage and Organization in the Brain
打开互动全文版(逐段中英对照 + 图/公式 + 论文问答)→如果我们要最终理解高等生物在感知识别、泛化、回忆和思维方面的能力,我们必须首先回答三个基本问题:1. 生物系统如何感知或检测物理世界的信息?2. 信息以何种形式存储或记忆?3. 存储或记忆中的信息如何影响识别和行为?这些问题中的第一个属于感官生理学范畴,也是目前唯一已获得相当理解的问题。本文主要关注第二和第三个问题,它们仍处于大量推测之中,而神经生理学目前提供的少数相关事实尚未整合成可接受的理论。关于第二个问题,存在两种对立观点。第一种观点认为,感觉信息的存储形式是编码表征或图像,在感觉刺激 与存储模式之间存在某种一一映射。根据这一假设,如果我们理解了神经系统的编码或“线路图”,原则上我们就能通过从刺激留下的“记忆痕迹”重建原始感觉模式,精确发现有机体记住了什么——就像我们可以冲洗照片底片,或翻译数字计算机“存储器”中的电荷模式。该假设因其简洁和易于理解而具有吸引力,并且围绕编码表征记忆这一思想已发展了大量理论脑模型(2, 3, 9, 14)。另一种方法源于英国经验主义传统,它大胆猜测刺激的图像可能从未真正被记录过,中枢神经系统仅仅作为一个复杂的交换网络,记忆以活动中心之间的新连接或通路形式存在。在这一立场的许多较新发展中(例如赫布的“细胞集合”和赫尔的“皮层预期目标反应”),与刺激关联的“反应”可能完全包含在 CNS 本身内部。在这种情况下,反应代表一种“观念”而非动作。该方法的重要特征是,永远不存在根据某种代码(允许后续重建)将刺激简单映射到记忆的过程。无论何种 信息被保留,都必须以对特定反应的偏好形式存储;即信息包含在连接或关联中,而非拓扑表征中。(在本文后续部分,“反应”一词应理解为有机体的任何可区分状态,可能涉及也可能不涉及外部可检测的肌肉活动。例如,中枢神经系统中某个细胞核的激活即可构成一个反应,根据这一定义。)对应于信息保留方法的这两种立场,关于第三个问题(存储信息如何影响当前活动)也存在两种假设。“编码记忆理论家”被迫得出结论:任何刺激的识别都涉及将存储内容与传入感觉模式进行匹配或系统比较,以确定当前刺激是否曾被见过,并确定有机体的适当反应。另一方面,经验主义传统的理论家实质上将第三个问题的答案与第二个问题的答案结合了起来:由于存储信息采取新连接或神经系统中传输通道的形式(或创建功能等同于新连接的条件),因此新刺激将利用这些已创建的新通路,自动激活适当反应,而无需单独的识别或辨识过程。本文将呈现的理论对这些问题的立场是经验主义的,或称“连接主义”。该理论是为一个假设的神经系统或机器(称为感知机)而发展的。感知机旨在说明智能系统的一些基本属性,而不深入涉足特定生物有机体所持有的特殊且常常未知的条件。感知机与生物系统之间的类比对于读者而言应当显而易见。近几十年来,符号逻辑、数字计算机和开关理论的发展使许多理论家认识到,神经元与构成计算机的简单开关单元
If we are eventually to understand the capability of higher organisms for perceptual recognition, generalization, recall, and thinking, we must first have answers to three fundamental questions: 1. How is information about the physical world sensed, or detected, by the biological system? 2. In what form is information stored, or remembered? 3. How does information contained in storage, or in memory, influence recognition and behavior? The first of these questions is in the province of sensory physiology, and is the only one for which appreciable understanding has been achieved. This article will be concerned primarily with the second and third questions, which are still subject to a vast amount of speculation, and where the few relevant facts currently supplied by neurophysiology have not yet been integrated into an acceptable theory. With regard to the second question, two alternative positions have been maintained. The first suggests that storage of sensory information is in the form of coded representations or images, with some sort of one-to-one mapping between the sensory stimulus and the stored pattern. According to this hypothesis, if one understood the code or "wiring diagr
如果我们要最终理解高等生物在感知识别、泛化、回忆和思维方面的能力,我们必须首先回答三个基本问题:1. 生物系统如何感知或检测物理世界的信息?2. 信息以何种形式存储或记忆?3. 存储或记忆中的信息如何影响识别和行为?这些问题中的第一个属于感官生理学范畴,也是目前唯一已获得相当理解的问题。本文主要关注第二和第三个问题,它们仍处于大量推测之中,而神经生理学目前提供的少数相关事实尚未整合成可接受的理论。关于第二个问题,存在两种对立观点。第一种观点认为,感觉信息的存储形式是编码表征或图像,在感觉刺激
If we are eventually to understand the capability of higher organisms for perceptual recognition, generalization, recall, and thinking, we must first have answers to three fundamental questions: 1. How is information about the physical world sensed, or detected, by the biological system? 2. In what form is information stored, or remembered? 3. How does information contained in storage, or in memory, influence recognition and behavior? The first of these questions is in the province of sensory physiology, and is the only one for which appreciable understanding has been achieved. This article will be concerned primarily with the second and third questions, which are still subject to a vast amount of speculation, and where the few relevant facts currently supplied by neurophysiology have not yet been integrated into an acceptable theory. With regard to the second question, two alternative positions have been maintained. The first suggests that storage of sensory information is in the form of coded representations or images, with some sort of one-to-one mapping between the sensory stimulus
与存储模式之间存在某种一一映射。根据这一假设,如果我们理解了神经系统的编码或“线路图”,原则上我们就能通过从刺激留下的“记忆痕迹”重建原始感觉模式,精确发现有机体记住了什么——就像我们可以冲洗照片底片,或翻译数字计算机“存储器”中的电荷模式。该假设因其简洁和易于理解而具有吸引力,并且围绕编码表征记忆这一思想已发展了大量理论脑模型(2, 3, 9, 14)。另一种方法源于英国经验主义传统,它大胆猜测刺激的图像可能从未真正被记录过,中枢神经系统仅仅作为一个复杂的交换网络,记忆以活动中心之间的新连接或通路形式存在。在这一立场的许多较新发展中(例如赫布的“细胞集合”和赫尔的“皮层预期目标反应”),与刺激关联的“反应”可能完全包含在 CNS 本身内部。在这种情况下,反应代表一种“观念”而非动作。该方法的重要特征是,永远不存在根据某种代码(允许后续重建)将刺激简单映射到记忆的过程。无论何种
and the stored pattern. According to this hypothesis, if one understood the code or "wiring diagram" of the nervous system, one should, in principle, be able to discover exactly what an organism remembers by reconstructing the original sensory patterns from the "memory traces" which they have left, much as we might develop a photographic negative, or translate the pattern of electrical charges in the "memory" of a digital computer. This hypothesis is appealing in its simplicity and ready intelligibility, and a large family of theoretical brain models has been developed around the idea of a coded, representational memory (2, 3, 9, 14). The alternative approach, which stems from the tradition of British empiricism, hazards the guess that the images of stimuli may never really be recorded at all, and that the central nervous system simply acts as an intricate switching network, where retention takes the form of new connections, or pathways, between centers of activity. In many of the more recent developments of this position (Hebb's "cell assembly," and Hull's "cortical anticipatory goal response," for example) the "responses" which are associated to stimuli may be entirely contained within the CNS itself. In this case the response represents an "idea" rather than an action. The important feature of this approach is that there is never any simple mapping of the stimulus into memory, according to some code which would permit its later reconstruction. Whatever in-
信息被保留,都必须以对特定反应的偏好形式存储;即信息包含在连接或关联中,而非拓扑表征中。(在本文后续部分,“反应”一词应理解为有机体的任何可区分状态,可能涉及也可能不涉及外部可检测的肌肉活动。例如,中枢神经系统中某个细胞核的激活即可构成一个反应,根据这一定义。)对应于信息保留方法的这两种立场,关于第三个问题(存储信息如何影响当前活动)也存在两种假设。“编码记忆理论家”被迫得出结论:任何刺激的识别都涉及将存储内容与传入感觉模式进行匹配或系统比较,以确定当前刺激是否曾被见过,并确定有机体的适当反应。另一方面,经验主义传统的理论家实质上将第三个问题的答案与第二个问题的答案结合了起来:由于存储信息采取新连接或神经系统中传输通道的形式(或创建功能等同于新连接的条件),因此新刺激将利用这些已创建的新通路,自动激活适当反应,而无需单独的识别或辨识过程。本文将呈现的理论对这些问题的立场是经验主义的,或称“连接主义”。该理论是为一个假设的神经系统或机器(称为感知机)而发展的。感知机旨在说明智能系统的一些基本属性,而不深入涉足特定生物有机体所持有的特殊且常常未知的条件。感知机与生物系统之间的类比对于读者而言应当显而易见。近几十年来,符号逻辑、数字计算机和开关理论的发展使许多理论家认识到,神经元与构成计算机的简单开关单元在功能上具有相似性,并提供了用此类元素表示高度复杂逻辑函数所需的分析方法。结果是出现了大量脑模型,它们不过是执行特定算法(代表“回忆”、刺激比较、变换和各种分析)以响应刺激序列的逻辑装置——例如 Rashevsky (14), McCulloch (10), McCulloch & Pitts (11), Culbertson (2), Kleene (8) 和 Minsky (13)。相对少数的理论家,如 Ashby (1) 和 von Neumann (17, 18),关注的是如何使包含许多随机连接的不完美神经网络可靠地执行那些可能由理想化线路图表示的功能。遗憾的是,符号逻辑和布尔代数语言不太适合此类研究。在仅能描述总体组织而精确结构未知的系统中,对事件进行数学分析需要一种合适的语言,这促使作者基于概率论而非符号逻辑构建当前模型。上述理论家主要关心的是,如何通过某种确定性的物理系统实现感知和回忆等功能,而非大脑实际如何做到。所提出的模型在若干重要方面(缺乏等势性、缺乏神经经济性、连接和同步要求过度特异、使细胞放电的刺激特异性不现实、假设了没有已知神经学关联的变量或功能特征等)均未能对应于生物系统。这种方法的支持者认为,一旦展示了如何让任何类型的物理系统感知和识别刺激或执行其他类脑功能,就只需对现有原则进行细化或修改,即可理解更真实神经系统的工作方式,并消除上述缺陷。另一方面,作者认为这些缺陷意味着仅仅对已有原则进行细化或改进永远无法解释生物智能;显然需要原则上的差异。这里将要总结的统计可分性理论(参见 15)似乎在原则上为所有这些困难提供了解决方案。那些更直接关注生物神经系统及其在自然环境中活动的理论家——赫布(7)、米尔纳(12)、埃克尔斯(4)、哈耶克(6)——通常表述不够精确,分析远非严格,因此常常难以评估他们所描述的系统是否能在真实的神经系统中实际工作,以及必要且充分的条件可能是什么。再次,缺乏可与网络分析师的布尔代数相媲美的分析语言是主要障碍之一。这一群体的贡献也许应被视为寻找和研究什么的建议,而非作为独立的、完整的理论系统。从这个角度看,就后续理论而言,最具启发性的工作是赫布和哈耶克的工作。由赫布(7)、哈耶克(6)、乌特利(16)和阿什比(1)特别阐述的立场是感知机理论的基础,可归纳为以下假设:1. 参与学习和识别的神经系统物理连接并非在每个有机体中完全相同。在出生时,最重要网络的构建很大程度上是随机的,仅受最少量的遗传约束。2. 原始连接细胞系统具有一定程度的可塑性;经过一段神经活动期后,刺激一组细胞导致另一组细胞产生响应的概率可能会因神经元本身的某些相对持久的改变而发生变化。3. 通过接触大量刺激样本,那些最“相似”(在某种必须根据特定物理系统定义的意义上)的刺激将倾向于形成通往同一组响应细胞的通路。那些明显“不相似”的刺激则倾向于发展到不同组响应细胞的连接。4. 正强化和/或负强化(或起此作用的刺激)的应用可能促进或阻碍当前正在进行的任何连接形成。5. 在此类系统中,相似性在神经系统的某个层次上表现为相似刺激倾向于激活同一组细胞。相似性并非特定形式或几何类刺激的必要属性,而是取决于感知系统的物理组织,这种组织通过与给定环境的相互作用而演化。系统的结构以及刺激环境的生态将影响并很大程度上决定感知世界被划分成的“事物”类别。感知机的组织 一个典型光感知机(以光学模式为刺激的感知机)的组织如图 1 所示。其组织规则如下:1. 刺激作用于感觉单元(S 点)的视网膜,在某些模型中假设这些单元以全或无方式响应,在另一些模型中则以与刺激强度成比例的脉冲幅度或频率响应。在本文考虑的模型中,将假设全或无响应。2. 冲动被传递到“投射区”(Ai)的一组关联细胞(A 单元)。在某些模型中,此投射区可省略,此时视网膜直接连接到关联区(An)。
formation is retained must somehow be stored as a preference for a particular response; i.e., the information is contained in connections or associations rather than topographic representations. (The term response, for the remainder of this presentation, should be understood to mean any distinguishable state of the organism, which may or may not involve externally detectable muscular activity. The activation of some nucleus of cells in the central nervous system, for example, can constitute a response, according to this definition.) Corresponding to these two positions on the method of information retention, there exist two hypotheses with regard to the third question, the manner in which stored information exerts its influence on current activity. The "coded memory theorists" are forced to conclude that recognition of any stimulus involves the matching or systematic comparison of the contents of storage with incoming sensory patterns, in order to determine whether the current stimulus has been seen before, and to determine the appropriate response from the organism. The theorists in the empiricist tradition, on the other hand, have essentially combined the answer to the third question with their answer to the second: since the stored information takes the form of new connections, or transmission channels in the nervous system (or the creation of conditions which are functionally equivalent to new connections), it follows that the new stimuli will make use of these new pathways which have been created, automatically activating the appropriate response without requiring any separate process for their recognition or identification. The theory to be presented here takes the empiricist, or "connectionist" position with regard to these questions. The theory has been developed for a hypothetical nervous system, or machine, called a perceptron. The perceptron is designed to illustrate some of the fundamental properties of intelligent systems in general, without becoming too deeply enmeshed in the special, and frequently unknown, conditions which hold for particular biological organisms. The analogy between the perceptron and biological systems should be readily apparent to the reader. During the last few decades, the development of symbolic logic, digital computers, and switching theory has impressed many theorists with the functional similarity between a neuron and the simple on-off units of which computers are constructed, and has provided the analytical methods necessary for representing highly complex logical functions in terms of such elements. The result has been a profusion of brain models which amount simply to logical contrivances for performing particular algorithms (representing "recall," stimulus comparison, transformation, and various kinds of analysis) in response to sequences of stimuli—e.g., Rashevsky (14), McCulloch (10), McCulloch & Pitts (11), Culbertson (2), Kleene (8), and Minsky (13). A relatively small number of theorists, like Ashby (1) and von Neumann (17, 18), have been concerned with the problems of how an imperfect neural network, containing many random connections, can be made to perform reliably those functions which might be represented by idealized wiring diagrams. Unfortunately, the language of symbolic logic and Boolean algebra is less well suited for such investigations. The need for a suitable language for the mathematical analysis of events in systems where only the gross organization can be characterized, and the 388 F. ROSENBLATT precise structure is unknown, has led the author to formulate the current model in terms of probability theory rather than symbolic logic. The theorists referred to above were chiefly concerned with the question of how such functions as perception and recall might be achieved by a deterministic physical system of any sort, rather than how this is actually done by the brain. The models which have been produced all fail in some important respects (absence of equipotentiality, lack of neuroeconomy, excessive specificity of connections and synchronization requirements, unrealistic specificity of stimuli sufficient for cell firing, postulation of variables or functional features with no known neurological correlates, etc.) to correspond to a biological system. The proponents of this line of approach have maintained that, once it has been shown how a physical system of any variety might be made to perceive and recognize stimuli, or perform other brainlike functions, it would require only a refinement or modification of existing principles to understand the working of a more realistic nervous system, and to eliminate the shortcomings mentioned above. The writer takes the position, on the other hand, that these shortcomings are such that a mere refinement or improvement of the principles already suggested can never account for biological intelligence; a difference in principle is clearly indicated. The theory of statistical separability (Cf. 15), which is to be summarized here, appears to offer a solution in principle to all of these difficulties. Those theorists—Hebb (7), Milner (12), Eccles (4), Hayek (6)—who have been more directly concerned with the biological nervous system and its activity in a natural environment, rather than with formally analogous machines, have generally been less exact in their formulations and far from rigorous in their analysis, so that it is frequently hard to assess whether or not the systems that they describe could actually work in a realistic nervous system, and what the necessary and sufficient conditions might be. Here again, the lack of an analytic language comparable in proficiency to the Boolean algebra of the network analysts has been one of the main obstacles. The contributions of this group should perhaps be considered as suggestions of what to look for and investigate, rather than as finished theoretical systems in their own right. Seen from this viewpoint, the most suggestive work, from the standpoint of the following theory, is that of Hebb and Hayek. The position, elaborated by Hebb (7), Hayek (6), Uttley (16), and Ashby (1), in particular, upon which the theory of the perceptron is based, can be summarized by the following assumptions: 1. The physical connections of the nervous system which are involved in learning and recognition are not identical from one organism to another. At birth, the construction of the most important networks is largely random, subject to a minimum number of genetic constraints. 2. The original system of connected cells is capable of a certain amount of plasticity; after a period of neural activity, the probability that a stimulus applied to one set of cells will cause a response in some other set is likely to change, due to some relatively long-lasting changes in the neurons themselves. 3. Through exposure to a large sample of stimuli, those which are most "similar" (in some sense which must be defined in terms of the particular physical system) will tend THE PERCEPTRON 389
投射区中的每个细胞从感觉点接收若干连接。将冲动传递到特定 A 单元的 S 点集称为该 A 单元的起源点。这些起源点对 A 单元的影响可以是兴奋性的或抑制性的。如果兴奋性和抑制性冲动强度的代数和等于或大于 A 单元的阈值(θ),则 A 单元放电,同样以全或无方式(或在某些模型中——此处不予考虑——以取决于所接收冲动净值的频率放电)。投射区中 A 单元的起源点倾向于围绕每个 A 单元的某个中心点聚集或局域化。随着到该 A 单元中心点的视网膜距离增加,起源点数量呈指数下降。(这种分布似乎得到生理学证据的支持,并在轮廓检测中起重要功能作用。)3. 投射区与关联区(An)之间的连接假设是随机的。即,An 集中的每个 A 单元从 AI 集中的起源点接收若干纤维,但这些起源点随机散布在整个投射区。除连接分布外,An 单元与 AI 单元相同,并在类似条件下响应。4. “响应”R1、R2、…
to form pathways to the same sets of responding cells. Those which are markedly "dissimilar" will tend to develop connections to different sets of responding cells. 4. The application of positive and/or negative reinforcement (or stimuli which serve this function) may facilitate or hinder whatever formation of connections is currently in progress. 5. Similarity, in such a system, is represented at some level of the nervous system by a tendency of similar stimuli to activate the same sets of cells. Similarity is not a necessary attribute of particular formal or geometrical classes of stimuli, but depends on the physical organization of the perceiving system, an organization which evolves through interaction with a given environment. The structure of the system, as well as the ecology of the stimulus-environment, will affect, and will largely determine, the classes of "things" into which the perceptual world is divided. THE ORGANIZATION OF A PERCEPTRON The organization of a typical photo-perceptron (a perceptron responding to optical patterns as stimuli) is shown in Fig. 1. The rules of its organization are as follows: 1. Stimuli impinge on a retina of sensory units (S-points), which are assumed to respond on an all-or-nothing basis, in some models, or with a pulse amplitude or frequency proportional to the stimulus intensity, in other models. In the models considered here, an all-or-nothing response will be assumed. 2. Impulses are transmitted to a set of association cells (A-units) in a "projection area" (Ai). This projection area may be omitted in some models, where the retina is connected directly to the association area (An).
The cells in the projection area each receive a number of connections from the sensory points. The set of S-points transmitting impulses to a particular A-unit will be called the origin points of that A-unit. These origin points may be either excitatory or inhibitory in their effect on the A-unit. If the algebraic sum of excitatory and inhibitory impulse intensities is equal to or greater than the threshold (θ) of the A-unit, then the A-unit fires, again on an all-or-nothing basis (or, in some models, which will not be considered here, with a frequency which depends on the net value of the impulses received). The origin points of the A-units in the projection area tend to be clustered or focalized, about some central point, corresponding to each A-unit. The number of origin points falls off exponentially as the retinal distance from the central point for the A-unit in question increases. (Such a distribution seems to be supported by physiological evidence, and serves an important functional purpose in contour detection.) 3. Between the projection area and the association area (An), connections are assumed to be random. That is, each A-unit in the An set receives some number of fibers from origin points in the AI set, but these origin points are scattered at random throughout the projection area. Apart from their connection distribution, the An units are identical with the AI units, and respond under similar conditions. 4. The "responses," Ri, R2, . . . ,
R 单元是细胞(或细胞集合),其反应方式与 A 单元大致相同。每个响应在 A 集合中拥有大量随机分布的起源点。将脉冲传输到特定响应的 A 单元集合称为该响应的源集合。(响应的源集合等同于它在 A 系统中的起源点集合。)图 1 中的箭头表示网络中的传输方向。注意,直到 An,所有连接都是前向的,没有反馈。当我们到达 An 和 R 单元之间的最后一组连接时,连接是双向建立的。在大多数感知机模型中,控制反馈连接的规则可以是以下两种备选之一:(a) 每个响应与其自身源集合中的细胞建立兴奋性反馈连接,或 (b) 每个响应与其自身源集合的补集建立抑制性反馈连接(即,它倾向于抑制任何不向其传输脉冲的关联细胞的活动)。第一条规则在解剖学上似乎更合理,因为 R 单元可能位于其各自源集合所在的同一皮层区域,
R-units are cells (or sets of cells) which respond in much the same fashion as the A-units. Each response has a typically large number of origin points located at random in the A set. The set of A-units transmitting impulses to a particular response will be called the source-set for that response. (The source-set of a response is identical to its set of origin points in the A-system.) The arrows in Fig. 1 indicate the direction of transmission through the network. Note that up to An all connections are forward, and there is no feedback. When we come to the last set of connections, between An and the R-units, connections are established in both directions. The rule governing feedback connections, in most models of the perceptron, can be either of the following alternatives: (a) Each response has excitatory feedback connections to the cells in its own source-set, or (b) Each response has inhibitory feedback connections to the complement of its own source-set (i.e., it tends to prohibit activity in any association cells which do not transmit to it). The first of these rules seems more plausible anatomically, since the R-units might be located in the same cortical area as their respective source-sets,
这使得 R 单元与适当源集合的 A 单元之间的相互兴奋极有可能发生。然而,备选规则(b)导致系统更易于分析,因此将假设用于这里要评估的大多数系统。图 2 显示了一个简化感知机的组织,这为统计可分性理论提供了便利的切入点。在为这个简化模型发展了理论之后,我们将能够更好地讨论图 1 中系统的优点。图 2 中显示的反馈连接是抑制性的,并且通向它们所源自响应的源集合的补集;因此,该系统按照上述规则 b 组织。这里显示的系统只有三级,第一级关联阶段已被消除。每个 A 单元在视网膜中有一组随机定位的起源点。这样的系统将基于刺激的重叠区域而非轮廓或外形的相似性来形成相似性概念。尽管这样的系统在许多辨别实验中处于劣势,但其能力仍然相当令人印象深刻,正如将很快证明的那样。图 2 中显示的系统只有两个响应,但显然可以包含的响应数量没有限制。以这种方式组织的系统中的响应是互斥的。如果 R1 发生,它将倾向于抑制 R2,并且也将抑制 R2 的源集合。同样,如果 R2 发生,它将倾向于抑制 R1。如果从一个源集合中的所有 A 单元接收的总脉冲比由另一个(对抗性)响应接收的脉冲更强或更频繁,则第一个响应将
making mutual excitation between the R-units and the A-units of the appropriate source-set highly probable. The alternative rule (b) leads to a more readily analyzed system, however, and will therefore be assumed for most of the systems to be evaluated here. Figure 2 shows the organization of a simplified perceptron, which affords a convenient entry into the theory of statistical separability. After the theory has been developed for this simplified model, we will be in a better position to discuss the advantages of the system in Fig. 1. The feedback connections shown in Fig. 2 are inhibitory, and go to the complement of the source-set for the response from which they originate; consequently, this system is organized according to Rule b, above. The system shown here has only three stages, the first association stage having been eliminated. Each A-unit has a set of randomly located origin points in the retina. Such a system will form similarity concepts on the basis of coincident areas of stimuli, rather than by the similarity of contours or outlines. While such a system is at a disadvantage in many discrimination experiments, its capability is still quite impressive, as will be demonstrated presently. The system shown in Fig. 2 has only two responses, but there is clearly no limit on the number that might be included. The responses in a system organized in this fashion are mutually exclusive. If R1 occurs, it will tend to inhibit R2, and will also inhibit the source-set for R2. Likewise, if R2 should occur, it will tend to inhibit R1. If the total impulse received from all the A-units in one source-set is stronger or more frequent than the impulse received by the alternative (antagonistic) response, then the first response will
倾向于获得对另一个响应的优势,并成为实际发生的响应。如果这样的系统要具备学习能力,那么必须能够修改 A 单元或其连接,使得一类刺激倾向于在 R1 源集合中比在 R2 源集合中诱发更强的脉冲,而另一类(不同)刺激倾向于在 R2 源集合中比在 R1 源集合中诱发更强的脉冲。假设每个 A 单元传递的脉冲可以用一个值 V 来表征,该值可以是幅度、频率、延迟或完成传输的概率。如果一个 A 单元具有高值,则其所有输出脉冲被认为比来自较低值 A 单元的脉冲更有效、更强力或更可能到达其终末突触。A 单元的值被认为是一个相当稳定的特征,可能取决于细胞和细胞膜的代谢状态,但并非绝对恒定。通常假设,活动期倾向于增加细胞的值,而值可能在(某些模型中)不活动时衰减。最有趣的模型是那些假设细胞竞争代谢物质的模型,更活跃的细胞以较不活跃的细胞为代价获得增益。在这样的系统中,如果没有活动,所有细胞将倾向于保持在相对恒定的状态,并且(无论活动与否)整个系统的净值
tend to gain an advantage over the other, and will be the one which occurs. If such a system is to be capable of learning, then it must be possible to modify the A-units or their connections in such a way that stimuli of one class will tend to evoke a stronger impulse in the R1 source-set than in the R2 source-set, while stimuli of another (dissimilar) class will tend to evoke a stronger impulse in the R2 source-set than in the R1 source-set. It will be assumed that the impulses delivered by each A-unit can be characterized by a value, V, which may be an amplitude, frequency, latency, or probability of completing transmission. If an A-unit has a high value, then all of its output impulses are considered to be more effective, more potent, or more likely to arrive at their endbulbs than impulses from an A-unit with a lower value. The value of an A-unit is considered to be a fairly stable characteristic, probably depending on the metabolic condition of the cell and the cell membrane, but it is not absolutely constant. It is assumed that, in general, periods of activity tend to increase a cell's value, while the value may decay (in some models) with inactivity. The most interesting models are those in which cells are assumed to compete for metabolic materials, the more active cells gaining at the expense of the less active cells. In such a system, if there is no activity, all cells will tend to remain in a relatively constant condition, and (regardless of activity) the net value of the system, taken in
始终保持不变。已经定量研究了三种类型的系统,它们在值动态上有所不同。它们的主要逻辑特征在表 1 中进行了比较。在 Alpha 系统中,活跃细胞仅每接收到一个脉冲就获得一个值的增量,并无限期地保持这个增益。在 Beta 系统中,每个源集合被允许一定的恒定增益率,增量按比例分配给源集合中活跃的细胞。在 Gamma 系统中,活跃细胞以其源集合中不活跃细胞为代价获得值增益,因此源集合的总值始终恒定。为了便于分析,通常将系统对刺激的响应分为两个阶段(图 3)。在主导阶段中,一定比例的 A 单元(图中以实心点表示)对刺激做出响应,但 R 单元仍处于非活动状态。该阶段是暂时的,并迅速进入后主导阶段,其中一个响应变得活跃,抑制其自身源集合的补集活动,
its entirety, will remain constant at all times. Three types of systems, which differ in their value dynamics, have been investigated quantitatively. Their principal logical features are compared in Table 1. In the alpha system, an active cell simply gains an increment of value for every impulse, and holds this gain indefinitely. In the beta system, each source-set is allowed a certain constant rate of gain, the increments being apportioned among the cells of the source-set in proportion to their activity. In the gamma system, active cells gain in value at the expense of the inactive cells of their source-set, so that the total value of a source-set is always constant. For purposes of analysis, it is convenient to distinguish two phases in the response of the system to a stimulus (Fig. 3). In the predominant phase, some proportion of A-units (represented by solid dots in the figure) responds to the stimulus, but the R-units are still inactive. This phase is transient, and quickly gives way to the postdominant phase, in which one of the responses becomes active, inhibiting activity in the com-
从而防止任何备选响应的发生。最初哪个响应变成主导是随机的,但如果 A 单元被强化(即,如果允许活跃单元获得值增益),那么当稍后再次呈现相同刺激时,同一响应将有更强的重复倾向,可以说学习已经发生。主导阶段分析:这里考虑的感知机将始终假设一个固定阈值。
plement of its own source-set, and thus preventing the occurrence of any alternative response. The response which happens to become dominant is initially random, but if the A-units are reinforced (i.e., if the active units are allowed to gain in value), then when the same stimulus is presented again at a later time, the same response will have a stronger tendency to recur, and learning can be said to have taken place. ANALYSIS OF THE PREDOMINANT PHASE The perceptrons considered here will always assume a fixed threshold,
\(P_a\),用于 A 单元的激活。这种系统称为固定阈值模型,与连续换能器模型形成对比,后者中 A 单元的响应是撞击刺激能量的某种连续函数。为了预测固定阈值感知机的学习曲线,发现有两个变量至关重要。它们的定义如下:
P_a, for the activation of the A-units. Such a system will be called a fixed-threshold model, in contrast to a continuous transducer model, where the response of the A-unit is some continuous function of the impinging stimulus energy. In order to predict the learning curves of a fixed-threshold perceptron, two variables have been found to be of primary importance. They are defined as follows:
\(P_a = \)给定大小的刺激所激活的 A 单元的期望比例,
P_a = the expected proportion of A-units activated by a stimulus of a given size,
\(P_C = \)一个对给定刺激\(S_i\)有响应的 A 单元也对另一个给定刺激\(S_2\)有响应的条件概率。可以证明(Rosenblatt, 1958),随着视网膜尺寸增大,S 点的数量(\(N_a\))很快就不再是一个重要参数,而\(P_a\)和\(P_c\)的值趋近于它们在具有无穷多点的视网膜上应有的值。因此,对于大视网膜,方程如下:
P_C = the conditional probability that an A-unit which responds to a given stimulus, S_i, will also respond to another given stimulus, S_2. It can be shown (Rosenblatt, 1958) that as the size of the retina is increased, the number of S-points (N_a) quickly ceases to be a significant parameter, and the values of P_a and P_c approach the value that they would have for a retina with infinitely many points. For a large retina, therefore, the equations are as follows:
\(R = \)刺激激活的 S 点的比例
R = proportion of S-points activated by the stimulus
\(x — \)每个 A 单元的兴奋性连接数量
x — number of excitatory connections to each A-unit
y = 每个 A 单元的抑制连接数量 θ = A 单元的阈值。(量 e 和 i 是 A 单元从刺激中接收到的兴奋性和抑制性分量。如果代数和 a = e + i 等于或大于θ,则假定该 A 单元响应。)
y = number of inhibitory connections to each A-unit θ = threshold of A-units. (The quantities e and i are the excitatory and inhibitory components of the excitation received by the A-unit from the stimulus. If the algebraic sum a = e + i is equal to or greater than θ, the A-unit is assumed to respond.)
L — 由第一个刺激 Si 照明的 S 点中未被(第二个刺激)照明的比例
L — proportion of the S-points illuminated by the first stimulus, Si, which are not illuminated by
G — 剩余 S 集(由第一个刺激留下)中被包含在第二个刺激(82)中的比例。量 R、L 和 G 指定了两个刺激及其视网膜重叠。
G — proportion of the residual S-set (left over from the first stimulus) which is included in the second stimulus (82). The quantities R, L, and G specify the two stimuli and their retinal overlap.
le 和 li 分别是在刺激 Si 被替换为 82 时 A 单元“丢失”的兴奋性和抑制性起源点的数量;ge 和
le and li are, respectively, the numbers of excitatory and inhibitory origin points "lost" by the A-unit when stimulus Si is replaced by 82; ge and
gi 是在刺激 Si 被替换为 82 时“获得”的兴奋性和抑制性起源点的数量。方程 2 中的求和是在指定界限之间进行的,受限于附加条件 e - i - I, + l{
gi are the numbers of excitatory and inhibitory origin points "gained" when stimulus Si is replaced by 82. The summations in Equation 2 are between the limits indicated, subject to the side condition e - i - I, + l{
Pa 的一些最重要特征在图 4 中展示,该图显示了 Pa 作为照明视网膜面积(R)的函数。注意可以通过增大阈值 6 来减小 Pa 的幅度,
Some of the most important characteristics of Pa are illustrated in Fig. 4, which shows Pa as a function of the retinal area illuminated (R). Note that Pa can be reduced in magnitude by either increasing the threshold, 6,
或者通过增加抑制性连接的比例(y)。比较图 4b 和 4c 显示,如果激发大约等于抑制,Pa 作为 R 的函数曲线被拉平,因此对于不同大小的刺激,Pa 变化很小。这一事实对于需要 Pa 接近最佳值才能正常工作的系统非常重要。
or by increasing the proportion of inhibitory connections (\(y\)). A comparison of Fig. 4b and 4c shows that if the excitation is about equal to the inhibition, the curves for \(P_a\) as a function of \(R\) are flattened out, so that there is little variation in \(P_a\) for stimuli of different sizes. This fact is of great importance for systems which require
Pa 需要接近最佳值才能正常工作。Pc 的行为如图 5 和图 6 所示。图 5 中的曲线可以与图 4 中 Pa 的曲线进行比较。注意,随着阈值增加,Pc 值的下降比 Pa 更为剧烈。Pc 也随着抑制性连接比例的增加而下降,就像 Pa 一样。图 5,即 394 F. ROSENBLATT。
\(P_a\) to be close to an optimum value in order to perform properly. The behavior of \(P_c\) is illustrated in Fig. 5 and 6. The curves in Fig. 5 can be compared with those for \(P_a\) in Fig. 4. Note that as the threshold is increased, there is an even sharper reduction in the value of \(P_c\) than was the case with \(P_a\). \(P_c\) also decreases as the proportion of inhibitory connections increases, as does \(P_a\). Fig. 5, which is 394 F. ROSENBLATT
图 4. P0 作为视网膜照明区域的函数。
FIG. 4. \(P_0\) as function of retinal area illuminated.
针对非重叠刺激计算,说明了即使刺激完全不相交且不照亮任何公共视网膜点,Pc 仍大于零。图 6 显示了刺激之间不同重叠量的影响。在所有情况下,Pc 的值
calculated for nonoverlapping stimuli, illustrates the fact that \(P_c\) remains greater than zero even when the stimuli are completely disjunct, and illuminate no retinal points in common. In Fig. 6, the effect of varying amounts of overlap between the stimuli is shown. In all cases, the value of \(P_c\)
随着刺激趋近于完全一致,Pc 趋向于 1。对于较小的刺激(虚线曲线),Pc 的值
goes to unity as the stimuli approach perfect identity. For smaller stimuli (broken line curves), the value of \(P_c\)
对于大刺激,该值较低。类似地,对于高阈值,该值小于低阈值。Pc 的最小值将等于
is lower than for large stimuli. Similarly, the value is less for high thresholds than for low thresholds. The minimum value of Pc will be equal to
$P_{cmin} = (1 - L)(1 - G)$。 (3) 在图 6 中,$P_{cmin}$对应于$θ = 10$的曲线。注意在这些条件下,A 单元对两个刺激都做出响应的概率
$P_{cmin} = (1 - L)(1 - G)$. (3) In Fig. 6, $P_{cmin}$ corresponds to the curve for $\theta = 10$. Note that under these conditions the probability that the A-unit responds to both stimuli
$(P_c)$实际上为零,除非刺激彼此非常接近。这种情况对判别学习可能有相当大的帮助。 感知机学习的数学分析 在主导阶段,感知机的响应中,一部分 A 单元(分散在整个系统中)对刺激做出反应,但很快让位于后主导响应,在此阶段活动局限于单个源集,其他源集被抑制。对于后主导阶段中“主导”响应的确定,已研究了两种可能的系统。第一种(均值判别系统,或μ系统)中,输入均值最大的响应首先做出反应,获得微弱优势,从而迅速成为主导。第二种(总和判别系统,或 S 系统)中,输入净值和最大的响应获得优势。在大多数情况下,响应均值的系统比响应总和的系统更具优势,因为均值受不同源集间$P_a$随机波动的影响较小。但在μ系统的情况下(见表 1),μ系统和 S 系统的性能变得相同。 我们已经指出,感知机通过学习或形成联想,是联想细胞活动导致值变化的结果。在评估这种学习时,可以考虑两类假设实验。第一类实验中,感知机暴露于一系列刺激模式(可能呈现于视网膜的随机位置),并“被迫”在每种情况下给出期望响应。(这种强制响应被认为是实验者的特权。在旨在评估试错学习的实验中,使用更复杂的感知机时,实验者并不强制系统以期望方式响应,而仅当响应正确时给予正强化,错误时给予负强化。)在评估这一“学习序列”期间发生的学习时,假设感知机“冻结”在当前状态,不再允许值变化,并以完全相同的方式再次呈现同一刺激序列,使刺激落在视网膜的相同位置上。感知机倾向于“正确”响应(即学习序列中先前被强化的响应)而非任何给定备选响应的概率称为$P_r$,即两个备选间正确选择响应的概率。第二类实验中,学习序列与之前完全相同地呈现,但评估感知机性能时不再使用之前展示过的刺激序列,而是呈现一个新序列,其中的刺激可能来自之前经历过的相同类别,但不一定完全相同。这一新测试序列假定由投射到随机视网膜位置的刺激组成,这些位置与学习序列选择的位置独立。测试序列的刺激在大小或旋转位置上也可能与之前经历过的刺激不同。在这种情况下,我们关心的是感知机对所代表刺激类别给出正确响应的概率,无论该特定刺激是否被见过。该概率称为$P_g$,即正确泛化的概率。与$P_r$一样,$P_g$实际上是发现偏向于正确响应而非任何其他备选的概率;每次只考虑一对响应,在一对中响应偏置正确并不意味着其他对中偏置不会偏向错误响应。正确响应优于所有备选的概率记为$P_R$或$P_G$。
$(P_c)$ is practically zero, except for stimuli which are quite close to identity. This condition can be of considerable help in discrimination learning. MATHEMATICAL ANALYSIS OF LEARNING IN THE PERCEPTRON The response of the perceptron in the predominant phase, where some fraction of the A-units (scattered throughout the system) responds to the stimulus, quickly gives way to the postdominant response, in which activity is limited to a single source-set, the other sets being suppressed. Two possible systems have been studied for the determination of the "dominant" response, in the postdominant phase. In one (the mean-discriminating system, or $\mu$-system), the response whose inputs have the greatest mean value responds first, gaining a slight advantage over the others, so that it quickly becomes dominant. In the second case (the sum-discriminating system, or S-system), the response whose inputs have the greatest net value gains an advantage. In most cases, systems which respond to mean values have an advantage over systems which respond to sums, since the means are less influenced by random variations in $P_a$ from one source-set to another. In the case of the $\mu$-system (see Table 1), however, the performance of the $\mu$-system and S-system become identical. We have indicated that the perceptron is expected to learn, or to form associations, as a result of the changes in value that occur as a result of the activity of the association cells. In evaluating this learning, one of two types of hypothetical experiments can be considered. In the first case, the perceptron is exposed to some series of stimulus patterns (which might be presented in random positions on the retina) and is "forced" to give the desired response in each case. (This forcing of responses is assumed to be a prerogative of the experimenter. In experiments intended to evaluate trial-and-error learning, with more sophisticated perceptrons, the experimenter does not force the system to respond in the desired fashion, but merely applies positive reinforcement when the response happens to be correct, and negative reinforcement when the response is wrong.) In evaluating the learning which has taken place during this "learning series," the perceptron is assumed to be "frozen" in its current condition, no further value changes being allowed, and the same series of stimuli is presented again in precisely the same fashion, so that the stimuli fall on identical positions on the retina. The probability that the perceptron will show a bias towards the "correct" response (the one which has been previously reinforced during the learning series) in preference to any given alternative response is called $P_r$, the probability of correct choice of response between two alternatives. In the second type of experiment, a learning series is presented exactly as before, but instead of evaluating the perceptron's performance using the same series of stimuli which were shown before, a new series is presented, in which stimuli may be drawn from the same classes that were previously experienced, but are not necessarily identical. This new test series is assumed to be composed of stimuli projected onto random retinal positions, which are chosen independently of the positions selected for the learning series. The stimuli of the test series may also differ in size or rotational position from the stimuli which were previously experienced. In this case, we are interested in the probability that the perceptron will give the correct response for the class of stimuli which is represented, regardless of whether the particular stimulus has been seen before or not. This probability is called $P_g$, the probability of correct generalization. As with $P_r$, $P_g$ is actually the probability that a bias will be found in favor of the proper response rather than any one alternative; only one pair of responses at a time is considered, and the fact that the response bias is correct in one pair does not mean that there may not be other pairs in which the bias favors the wrong response. The probability that the correct response will be preferred over all alternatives is designated $P_R$ or $P_G$.
这一新测试序列假定由投射到随机视网膜位置的刺激组成,这些位置与学习序列选择的位置独立。测试序列的刺激在大小或旋转位置上也可能与之前经历过的刺激不同。在这种情况下,我们关心的是感知机对所代表刺激类别给出正确响应的概率,无论该特定刺激是否被见过。该概率称为$P_g$,即正确泛化的概率。与$P_r$一样,$P_g$实际上是发现偏向于正确响应而非任何其他备选的概率;每次只考虑一对响应,在一对中响应偏置正确并不意味着其他对中偏置不会偏向错误响应。正确响应优于所有备选的概率记为$P_R$或$P_G$。
This new test series is assumed to be composed of stimuli projected onto random retinal positions, which are chosen independently of the positions selected for the learning series. The stimuli of the test series may also differ in size or rotational position from the stimuli which were previously experienced. In this case, we are interested in the probability that the perceptron will give the correct response for the class of stimuli which is represented, regardless of whether the particular stimulus has been seen before or not. This probability is called $P_g$, the probability of correct generalization. As with $P_r$, $P_g$ is actually the probability that a bias will be found in favor of the proper response rather than any one alternative; only one pair of responses at a time is considered, and the fact that the response bias is correct in one pair does not mean that there may not be other pairs in which the bias favors the wrong response. The probability that the correct response will be preferred over all alternatives is designated $P_R$ or $P_G$.
is lower than for large stimuli. Simi-larly, the value is less for high thresh-olds than for low thresholds. The minimum value of Pc will be equal to
在所有研究的情况中,只要代入适当的常数,一个通用的方程就可以很好地近似 Pr 和 Pg。该方程的形式为:\P = P(N_{ar} > 0) \cdot \Phi(Z)\ (4),其中
In all cases investigated, a single general equation gives a close approximation to Pr and Pg, if the appropriate constants are substituted. This equation is of the form: \P = P(N_{ar} > 0) \cdot \Phi(Z)\ (4) where
\Phi(Z)\ 是从负无穷到 Z 的正态曲线积分。
\Phi(Z) = \text{normal curve integral from } -\infty \text{ to } Z\
是所考虑的备选反应。方程 4 是在学习期间对两个反应各呈现 n, r 个刺激后,R_j 被偏好于 R_a 的概率。N 是每个源集合中“有效”A 单元的数量;即,任一源集合中不与两个反应共同连接的 A 单元的数量。那些共同连接的单元对价值平衡的两边贡献相等,因此不影响对一个反应或另一个反应的净偏向。
is the alternative response under consideration. Equation 4 is the probability that R_j will be preferred over R_a after n, r stimuli have been shown for each of the two responses, during the learning period. N is the number of "effective" A-units in each source-set; that is, the number of A-units in either source-set which are not connected in common to both responses. Those units which are connected in common contribute equally to both sides of the value balance, and consequently do not affect the net bias towards one response or the other.
N_{ar} 是源集合中对测试刺激 St 作出响应的活跃单元数量。P(N_{ar} > 0) 是至少一个 N_e
N_{ar} is the number of active units in a source-set, which respond to the test stimulus, St. P(N_{ar} > 0) is the probability that at least one of the N_e
有效单元(按惯例指定为 R_i 反应)将被测试刺激 St 激活。
effective units in the source-set of the correct response (designated, by convention, as the R_i response) will be activated by the test stimulus, St.
对于 Pg,常数 \c_2\ 始终等于零,其他三个常数与 Pr 相同。
In the case of Pg, the constant \c_2\ is always equal to zero, the other three constants being the same as for Pr.
四个常数的值取决于物理神经网络(感知机)的参数以及刺激环境的组织。最容易分析的情况是向感知机展示来自“理想环境”的刺激,该环境由随机放置的光点组成,不试图根据内在相似性对刺激进行分类。因此,在一个典型的学习实验中,我们可能向感知机展示由随机照亮的视网膜点组成的 1,000 个刺激,并任意强化 \R_1\ 作为前 500 个刺激的“正确”响应,强化 \R_2\ 作为剩余 500 个的响应。这种环境之所以“理想”,只是我们像物理学中谈论理想气体那样;它是一个便于分析的人造物,并不会导致感知机的最佳性能。在理想环境情况下,常数 \c_1\ 始终等于零,因此,在 THE PERCEPTRON 397
The values of the four constants depend on the parameters of the physical nerve net (the perceptron) and also on the organization of the stimulus environment. The simplest cases to analyze are those in which the perceptron is shown stimuli drawn from an "ideal environment," consisting of randomly placed points of illumination, where there is no attempt to classify stimuli according to intrinsic similarity. Thus, in a typical learning experiment, we might show the perceptron 1,000 stimuli made up of random collections of illuminated retinal points, and we might arbitrarily reinforce \R_1\ as the "correct" response for the first 500 of these, and \R_2\ for the remaining 500. This environment is "ideal" only in the sense that we speak of an ideal gas in physics; it is a convenient artifact for purposes of analysis, and does not lead to the best performance from the perceptron. In the ideal environment situation, the constant \c_1\ is always equal to zero, so that, in the THE PERCEPTRON 397
对于 Pg 的情况(其中 \c_2\ 也为零),\Z\ 的值将为零,并且 Pg 永远无法优于随机期望的 0.5。然而,在这些条件下对 Pr 的评估却揭示了 α、β 和 γ 系统之间的一些有趣差异(表 1)。首先考虑 α 系统,它是三者中动力学最简单的。在该系统中,每当一个 A 单元活跃一个单位时间,它就获得一个单位价值。我们最初假设一个实验,其中 \N_r\(每个响应关联的刺激数量)对所有响应是常数。在这种情况下,对于求和系统,
case of Pg (where \c_2\ is also zero), the value of \Z\ will be zero, and Pg can never be any better than the random expectation of 0.5. The evaluation of Pr for these conditions, however, throws some interesting light on the differences between the alpha, beta, and gamma systems (Table 1). First consider the alpha system, which has the simplest dynamics of the three. In this system, whenever an A-unit is active for one unit of time, it gains one unit of value. We will assume an experiment, initially, in which \N_r\ (the number of stimuli associated to each response) is constant for all responses. In this case, for the sum system,
其中 \w = \frac{1}{N_R}\ 是连接到每个 A 单元的响应比例。如果源集合是不相交的,则 \w = \frac{1}{N_R}\,其中 \N_R\ 是系统中的响应数量。对于 γ 系统,
where \w = \frac{1}{N_R}\ is the fraction of responses connected to each A-unit. If the source-sets are disjunct, \w = \frac{1}{N_R}\, where \N_R\ is the number of responses in the system. For the γ-system,
In the case of Pg, the constant c 2 is always equal to zero, the other three constants being the same as for Pr.
将 \(c_3\) 减少到零使得 ^-系统比 S-系统具有明显优势。图 7 和图 8 比较了这些系统的典型学习曲线。图 9 显示了 \(P_a\) 变化对系统性能的影响。如果 \(n\) 和 \(r\) 不是固定的,而是作为随机变量处理,那么每个响应对应的刺激数量从某个分布中单独抽取,则性
The reduction of \(c_3\) to zero gives the ^-system a definite advantage over the S-system. Typical learning curves for these systems are compared in Fig. 7 and 8. Figure 9 shows the effect of variations in \(P_a\) upon the performance of the system. If \(n\) and \(r\), instead of being fixed, are treated as random variables, so that the number of stimuli associated to each response is drawn separately from some distribution, then the per-
能(performance)的 a-系统比上述方程所指示的要差得多。在这些条件下,//-系统的常数是
formance of the a-system is considerably poorer than the above equations indicate. Under these conditions, the constants for the //-system are
q 是 \(n_r\) 与 \(n_s\) 的比值,NR 是系统中的响应数量。
q = ratio of \(n_r\) to \(n_s\), NR = number of responses in the system
对于这个方程(以及其他将 \(n_r\) 视为随机变量的方程),有必要将 \(n_r\) 定义为
For this equation (and any others in which \(n_r\) is treated as a random variable), it is necessary to define \(n_r\)
方程 4 中该变量在所有响应集合上的期望值。对于 \(\beta\) 系统,性能的缺陷更大,这是因为无论系统发生什么,净值都会持续增长。刺激激活的子集的大净值倾向于放大小的统计差异,导致不可靠的性能。这种情况下的常数(再次对于 \(\mu\) 系统)是
in Equation 4 as the expected value of this variable, over the set of all responses. For the \(\beta\)-system, there is an even greater deficit in performance, due to the fact that the net value continues to grow regardless of what happens to the system. The large net values of the subsets activated by a stimulus tend to amplify small statistical differences, causing an unreliable performance. The constants in this case (again for the \(\mu\)-system) are
在 α 和 β 系统中,求和判别模型的性能将比均值判别模型更差。然而,在 γ 系统中,可以证明 \(P_e(s) = P_K / O\);即,使用 S 系统还是 \(\gamma\) 系统对性能没有影响。此外,对于 \(\gamma\)-
In both the alpha and beta systems, performance will be poorer for the sum-discriminating model than for the mean-discriminating case. In the gamma-system, however, it can be shown that \(P_e(s) = P_K / O\); i.e., it makes no difference in performance whether the S-system or \(\gamma\)-system is used. Moreover, the constants for the \(\gamma\)-
系统,具有变量 \(n_{sr}\),其常数与 α 系统相同。感知机 399
system, with variable \(n_{sr}\), are identical to the constants for the alpha system. THE PERCEPTRON 399
图 9. P(T) 作为 P* 的函数。(对于 n, T = 1,000, u = 0。假设理想环境。)
FIG. 9. P(T) as function of P*. (For n, T = 1,000, u = 0. Ideal environment assumed.)
项,固定 n_sr(方程 6),展示了 7-系统的优势。三个系统的性能在图 10 中进行了比较,该图清楚地表明……现在让我们用“差异环境”模型替换“理想环境”
Term, with n_sr fixed (Equation 6), demonstrates the advantage of the 7-system. The performance of the three systems is compared in Fig. 10, which clearly shows... Let us now replace the "ideal environment"
图 10. α、β和γ系统在不同 n_sr 下的比较(N_R = 100, n_sr = 0.5, N_A = 10,000, P_a = 0.07, w = 0.2)。400 F. ROSENBLATT 假设,用“差异环境”模型替换“理想环境”,在该模型中存在若干可区分的刺激类别(如正方形、圆形和三角形,或字母表中的字母)。如果我们设计一个实验,其中与每个响应相关的刺激来自不同类别,那么感知机的学习曲线将发生显著改变。最重要的区别在于常数 c₁(Z 分子中 n_sr 的系数)不再为零,因此方程 4 现在有一个非随机渐近线。此外,在 P_r(正确泛化概率)的形式中,c₂ = 0,量 Z 保持大于零,而 P_a 实际上趋近于与 P_r 相同的渐近线。因此,对于每个刺激类别经过无限经验后,感知机性能的方程对 P_T 和 P_g 是相同的:
FIG. 10. Comparison of α, β, and γ systems, for variable n_sr, (N_R = 100, n_sr = 0.5, N_A = 10,000, P_a = 0.07, w = 0.2). 400 F. ROSENBLATT environment" assumptions with a model for a "differentiated environment," in which several distinguishable classes of stimuli are present (such as squares, circles, and triangles, or the letters of the alphabet). If we then design an experiment in which the stimuli associated to each response are drawn from a different class, then the learning curves of the perceptron are drastically altered. The most important difference is that the constant c_1 (the coefficient of n_sr in the numerator of Z) is no longer equal to zero, so that Equation 4 now has a nonrandom asymptote. Moreover, in the form for P_r (the probability of correct generalization), where c_2 = 0, the quantity Z remains greater than zero, and P_a actually approaches the same asymptote as P_r. Thus the equation for the perceptron's performance after infinite experience with each class of stimuli is identical for P_T and P_g:
这意味着在极限情况下,感知机之前是否见过某个特定的测试刺激并无区别;如果刺激来自差异环境,性能在两种情况下同样好。
This means that in the limit it makes no difference whether the perceptron has seen a particular test stimulus before or not; if the stimuli are drawn from a differentiated environment, the performance will be equally good in either case.
FIG. 9. P,tf) as function of P*. (For n, T — 1,000, u, = 0. Ideal environment assumed.)
为了评估系统在差异环境中的性能,有必要定义量 \(P_{C\alpha\beta}\)。该量被解释为……的期望值。
In order to evaluate the performance of the system in a differentiated environment, it is necessary to define the quantity \(P_{C\alpha\beta}\). This quantity is interpreted as the expected value of
\(P_C\) 是指从类别 \(\alpha\) 和 \(\beta\) 中随机抽取的刺激对之间的值。特别地,\(P_{C\alpha\alpha}\) 是同类成员之间的 \(P_C\) 期望值,\(P_{C\alpha\beta}\) 是……的期望值。
\(P_C\) between pairs of stimuli drawn at random from classes \(\alpha\) and \(\beta\). In particular, \(P_{C\alpha\alpha}\) is the expected value of \(P_C\) between members of the same class, and \(P_{C\alpha\beta}\) is the expected value of
\(P_C\) 是指从类别 1 抽取的 \(S_1\) 刺激与从类别 2 抽取的 \(S_2\) 刺激之间的值。\(P_{C1x}\) 是……的期望值。
\(P_C\) between an \(S_1\) stimulus drawn from Class 1 and an \(S_2\) stimulus drawn from Class 2. \(P_{C1x}\) is the expected value of
如果 \(P_{C11} > P_\alpha > P_{C12}\),则感知机的极限性能 (\(P_{Boo}\)) 将优于随机水平,并且最终应出现对类别 1 成员的正确“泛化响应” \(R_I\) 的学习。如果不满足上述不等式,则可能不会出现优于随机水平的改进,而是可能出现类别 2 响应。可以证明 (IS) 对于大多数简单的几何形状(我们通常认为“相似”),如果系统参数选择得当,所需的不等式可以满足。对于总和判别版本的 \(\alpha\)-感知机,在差异环境中,所有响应的 \(n\) 和 \(r\) 固定时,\(P_r\) 方程将具有以下四个系数的表达式:
\(P_C\) between members of Class 1 and stimuli drawn at random from all other classes in the environment. If \(P_{C11} > P_\alpha > P_{C12}\), the limiting performance of the perceptron (\(P_{Boo}\)) will be better than chance, and learning of some response, \(R_I\), as the proper "generalization response" for members of Class 1 should eventually occur. If the above inequality is not met, then improvement over chance performance may not occur, and the Class 2 response is likely to occur instead. It can be shown (IS) that for most simple geometrical forms, which we ordinarily regard as "similar," the required inequality can be met, if the parameters of the system are properly chosen. The equation for \(P_r\), for the sum-discriminating version of an \(\alpha\)-perceptron, in a differentiated environment where \(n\) and \(r\) are fixed for all responses, will have the following expressions for the four coefficients:
ir) 和 \(\sigma^2(P_{C\alpha\beta})\) 表示 \(P_{Cir}\) 和 \(P_{C\alpha\beta}\) 在可能测试刺激集合 \(S_t\) 上的方差,以及《感知机》401 页。
ir) and \(\sigma^2(P_{C\alpha\beta})\) represent the variance of \(P_{Cir}\) and \(P_{C\alpha\beta}\) measured over the set of possible test stimuli, \(S_t\), and THE PERCEPTRON 401
σ_J 和σ^2(P_{clx})代表 P_{cir}和 P_{clx}在所有 A 单元集合上测得的方差,e = P_{clr} P_{clx}的协方差,假定可忽略。这些表达式中出现的方差迄今未能精确分析,可作为待定经验变量,根据所考虑的刺激类别确定。若σ设为变量期望值的一半,则在每种情况下可得保守估计。当给定类别的刺激形状相同且均匀分布在视网膜上时,下标 s 的方差等于零。Paw 将由同一组系数表示,除了 c_2 通常为零。对于均值判别系统,系数为:
σ_J and σ^2(P_{clx}) represent the variance of P_{cir} and P_{clx} measured over the set of all A-units, e = covariance of P_{clr} P_{clx}, which is assumed to be negligible. The variances which appear in these expressions have not yielded, thus far, to a precise analysis, and can be treated as empirical variables to be determined for the classes of stimuli in question. If the sigma is set equal to half the expected value of the variable, in each case, a conservative estimate can be obtained. When the stimuli of a given class are all of the same shape, and uniformly distributed over the retina, the subscript s variances are equal to zero. Paw will be represented by the same set of coefficients, except for c_2, which is equal to zero, as usual. For the mean-discriminating system, the coefficients are:
此处省略了一些视为可忽略的协方差项。图 11 展示了均值判别系统在差异化环境模型下的典型学习曲线集。这些参数基于对某个
Some covariance terms, which are considered negligible, have been omitted here. A set of typical learning curves for the differentiated environment model is shown in Fig. 11, for the mean-discriminating system. The parameters are based on measurements for a
F. ROSENBLATT 方-圆辨别问题。注意 Pr 和 Pg 的曲线
F. ROSENBLATT square-circle discrimination problem. Note that the curves for Pr and Pg
均如预期趋近同一渐近线。这些渐近线的值可通过将适当系数代入方程 9 得到。随着系统中关联细胞数量的增加,渐近学习极限迅速趋近于 1,因此对于包含数千个细胞的系统,在如此简单的问题上,性能误差应可忽略。随着系统中响应数量的增加,若每个响应与其他所有备选响应互斥,则性能逐步下降。避免这一退化的一种方法(详见 Rosenblatt, 15)是通过响应的二元编码。在此情况下,不再用 100 个不同且互斥的响应表示 100 种刺激模式,而是找到有限数量的判别特征,每个特征可独立识别存在与否,从而可用一对互斥响应表示。给定一组理想的二元特征(如亮/暗、高/矮、直/曲等),仅需七个响应对的恰当配置即可区分 100 种刺激类别。在对系统的进一步改进中,单个响应可通过其活跃或不活跃状态表示每个二元特征的存在与否。此类编码的效率取决于能用于区分刺激的独立可识别“标志”的数量。若刺激只能整体识别,无法进行此类分析,则最终需要单独的二元响应对(即比特)来表示每一刺激类别的存在与否(如“狗”或“非狗”),如此并未比所有响应互斥的系统带来任何增益。
both approach the same asymptotes, as predicted. The values of these asymptotes can be obtained by substituting the proper coefficients in Equation 9. As the number of association cells in the system increases, the asymptotic learning limit rapidly approaches unity, so that for a system of several thousand cells, the errors in performance should be negligible on a problem as simple as the one illustrated here. As the number of responses in the system increases, the performance becomes progressively poorer, if every response is made mutually exclusive of all alternatives. One method of avoiding this deterioration (described in detail in Rosenblatt, 15) is through the binary coding of responses. In this case, instead of representing 100 different stimulus patterns by 100 distinct, mutually exclusive responses, a limited number of discriminating features is found, each of which can be independently recognized as being present or absent, and consequently can be represented by a single pair of mutually exclusive responses. Given an ideal set of binary characteristics (such as dark, light; tall, short; straight, curved; etc.), 100 stimulus classes could be distinguished by the proper configuration of only seven response pairs. In a further modification of the system, a single response is capable of denoting by its activity or inactivity the presence or absence of each binary characteristic. The efficiency of such coding depends on the number of independently recognizable "earmarks" that can be found to differentiate stimuli. If the stimulus can be identified only in its entirety and is not amenable to such analysis, then ultimately a separate binary response pair, or bit, is required to denote the presence or absence of each stimulus class (e.g., "dog" or "not dog"), and nothing has been gained over a system where all responses are mutually exclusive.
迄今为止所分析的所有系统中,活跃 A 单元因强化或经验而获得的价值增量始终为正数,即活跃单元总能增强其激活所连接响应的能力。在γ-系统中,确实有些单元会失去价值,但这些单元总是非活跃单元,而活跃单元则按其活动速率成比例地获得价值。在二值系统中,有两种可能的强化(正强化和负强化),活跃单元可能根据系统瞬时状态而获得或失去价值。如果可以通过施加外部刺激来控制正强化和负强化,那么它们本质上等同于“奖励”和“惩罚”,实验者可以以此方式使用。在这些条件下,感知机似乎能够进行试错学习。然而,二值系统不一定需要应用奖励和惩罚。如果二进制编码的响应系统组织为:每个“比特”或所学习的刺激特征由一个单一响应或响应对表示,当响应为“开”时对其自身源集进行正反馈,当响应为“关”时进行负反馈(意味着活跃 A 单元将失去而非获得价值),那么该系统在其特性上仍然是二值的。这样的二值系统(THE PERCEPTRON 403)
In all of the systems analyzed up to this point, the increments of value gained by an active A-unit, as a result of reinforcement or experience, have always been positive, in the sense that an active unit has always gained in its power to activate the responses to which it is connected. In the gamma-system, it is true that some units lose value, but these are always the inactive units, the active ones gaining in proportion to their rate of activity. In a bivalent system, two types of reinforcement are possible (positive and negative), and an active unit may either gain or lose in value, depending on the momentary state of affairs in the system. If the positive and negative reinforcement can be controlled by the application of external stimuli, they become essentially equivalent to "reward" and "punishment," and can be used in this sense by the experimenter. Under these conditions, a perceptron appears to be capable of trial-and-error learning. A bivalent system need not necessarily involve the application of reward and punishment, however. If a binary-coded response system is so organized that there is a single response or response-pair to represent each "bit," or stimulus characteristic that is learned, with positive feedback to its own source-set if the response is "on," and negative feedback (in the sense that active A-units will lose rather than gain in value) if the response is "off," then the system is still bivalent in its characteristics. Such a bivalent THE PERCEPTRON 403
系统在减少困扰其他系统的某些偏差效应(因关联刺激的规模或频率较大而导致对错误响应的偏好)方面特别高效。已经考虑了多种形式的二值系统(15,第 VII 章)。其中最高效的具有以下逻辑特征:如果系统处于正强化状态,则对“开”响应源集中的所有活跃 A 单元添加正\(Delta V\),而对“关”响应源集中的活跃 A 单元添加负\(Delta V\)。如果系统当前处于负强化状态,则对“开”响应源集中的所有活跃 A 单元添加负\(Delta V\),而对“关”响应源集中的活跃 A 单元添加正\(Delta V\)。如果源集是不相交的(这是系统正常工作的必要条件),则对于\(mu\)情况,二值γ-系统的方程与单值α-系统具有相同的系数(方程 11)。该系统的性能曲线如图 12 所示,其中绘制了系统可达到的渐近泛化概率,使用的刺激参数与图 11 相同。这是 n-比特响应模式中所有比特都正确的概率。显然,如果多数正确响应足以正确识别刺激,则性能将优于这些曲线所示。在一种利用更合理的生物学假设的二值系统形式中,A 单元对连接响应的影响可以是兴奋性或抑制性的。该系统中的正\(Delta V\)对应于兴奋性单元的增长,而负\(Delta V\)对应于抑制性单元的增长。这种系统的性能与上述系统相似,但可以证明其效率较低。与图 12 所示的二值系统类似的系统已在康奈尔航空实验室的 IBM 704 计算机上的一系列实验中得到详细模拟。结果验证了理论的所有主要预测,并将另行报告。改进的感知机与自发组织 前面对感知机性能的定量分析未考虑时间作为刺激维度。不具备时间模式识别能力的感知机被称为“瞬时刺激感知机”。可以证明(15),只要刺激留下某种暂时持续的痕迹(例如,改变阈值),相同的统计可分性原理将使感知机能够区分速度、声音序列等。
system is particularly efficient in reducing some of the bias effects (preference for the wrong response due to greater size or frequency of its associated stimuli) which plague the alternative systems. Several forms of bivalent systems have been considered (15, Chap. VII). The most efficient of these has the following logical characteristics. If the system is under a state of positive reinforcement, then a positive \(Delta V\) is added to the values of all active A-units in the source-sets of "on" responses, while a negative \(Delta V\) is added to the active units in the source-sets of "off" responses. If the system is currently under negative reinforcement, then a negative \(Delta V\) is added to all active units in the source-set of an "on" response, and a positive \(Delta V\) is added to active units in an "off" source-set. If the source-sets are disjunct (which is essential for this system to work properly), the equation for a bivalent \(gamma\)-system has the same coefficients as the monovalent \(alpha\)-system, for the \(mu\)-case (Equation 11). The performance curves for this system are shown in Fig. 12, where the asymptotic generalization probability attainable by the system is plotted for the same stimulus parameters that were used in Fig. 11. This is the probability that all bits in an n-bit response pattern will be correct. Clearly, if a majority of correct responses is sufficient to identify a stimulus correctly, the performance will be better than these curves indicate. In a form of bivalent system which utilizes more plausible biological assumptions, A-units may be either excitatory or inhibitory in their effect on connected responses. A positive \(Delta V\) in this system corresponds to the incrementing of an excitatory unit, while a negative \(Delta V\) corresponds to the incrementing of an inhibitory unit. Such a system performs similarly to the one considered above, but can be shown to be less efficient. Bivalent systems similar to those illustrated in Fig. 12 have been simulated in detail in a series of experiments with the IBM 704 computer at the Cornell Aeronautical Laboratory. The results have borne out the theory in all of its main predictions, and will be reported separately at a later time. IMPROVED PERCEPTRONS AND SPONTANEOUS ORGANIZATION The quantitative analysis of perceptron performance in the preceding sections has omitted any consideration of time as a stimulus dimension. A perceptron which has no capability for temporal pattern recognition is referred to as a "momentary stimulus perceptron." It can be shown (15) that the same principles of statistical separability will permit the perceptron to distinguish velocities, sound sequences, etc., provided the stimuli leave some temporarily persistent trace, such as an altered threshold,
F. ROSENBLATT 使得 A 系统中的活动在时间 t 时在一定程度上依赖于时间 t−1 时的活动。另外还假设 A 单元的起点是完全随机的。可以证明,通过适当地组织起点(如图 1 所示的投影区域起点中空间分布受到约束),A 单元将对轮廓位置特别敏感,从而提高性能。在最近的一项发展中(我们希望在不久的将来详细报告),已经证明如果允许 A 单元的值以与其大小成比例的速率衰减,就会出现一个引人注目的新特性:感知机能够进行“自发”概念形成。也就是说,如果系统暴露于两个“不相似”类别的随机刺激序列中,并且其所有响应都被自动强化而不考虑它们是“正确”还是“错误”,那么系统将趋向于一个稳定的终态,其中(对于每个二值响应)对于一类刺激的成员响应为“1”,对于另一类刺激的成员响应为“0”;即感知机将自发地识别两个类别之间的差异。这一现象已在 704 计算机的模拟实验中得到成功演示。即使只有一个逻辑层的 A 单元和响应单元,感知机也可以被证明在选择性回忆和选择性注意领域具有许多有趣的性质。这些性质通常取决于不同响应的源集的交集,并在其他地方详细讨论(15)。通过结合音频和照片输入,可以将声音或听觉“名称”与视觉对象关联起来,并使感知机执行诸如“说出左边的物体”或“说出该刺激的颜色”等选择性响应。此时可能会提出一个问题:感知机的能力究竟止步于何处?我们已经看到,所描述的系统足以进行模式识别、联想学习以及选择性注意和选择性回忆所需的认知集。该系统似乎具有时间模式识别以及空间识别的潜力,涉及任何感觉模态或模态组合。可以证明,通过适当的强化,它能够进行试错学习,并且如果其自身响应通过感觉通道反馈,它可以学习发出有序的响应序列。这是否意味着感知机无需进一步原理上的修改就能具备诸如人类言语、交流和思维等高级功能?实际上,感知机能力的限制似乎在于相对判断和关系抽象领域。在其“符号行为”中,感知机表现出与 Goldstein 的脑损伤患者(5)的某些惊人相似之处。可以学习对明确的具体刺激的响应,即使正确的响应需要识别多个同时出现的限定条件(例如,如果刺激在左边则说出颜色,如果在右边则说出形状)。然而,一旦响应要求识别刺激之间的关系(如“说出正方形左边的物体”或“指出圆圈之前出现的模式”),感知机就变得极其困难(THE PERCEPTRON 405)
F. ROSENBLATT which causes the activity in the A-system at time t to depend to some degree on the activity at time t − 1. It has also been assumed that the origin points of A-units are completely random. It can be shown that by a suitable organization of origin points, in which the spatial distribution is constrained (as in the projection area origins shown in Fig. 1), the A-units will become particularly sensitive to the location of contours, and performance will be improved. In a recent development, which we hope to report in detail in the near future, it has been proven that if the values of the A-units are allowed to decay at a rate proportional to their magnitude, a striking new property emerges: the perceptron becomes capable of "spontaneous" concept formation. That is to say, if the system is exposed to a random series of stimuli from two "dissimilar" classes, and all of its responses are automatically reinforced without any regard to whether they are "right" or "wrong," the system will tend towards a stable terminal condition in which (for each binary response) the response will be "1" for members of one stimulus class, and "0" for members of the other class; i.e., the perceptron will spontaneously recognize the difference between the two classes. This phenomenon has been successfully demonstrated in simulation experiments, with the 704 computer. A perceptron, even with a single logical level of A-units and response units, can be shown to have a number of interesting properties in the field of selective recall and selective attention. These properties generally depend on the intersection of the source sets for different responses, and are elsewhere discussed in detail (15). By combining audio and photo inputs, it is possible to associate sounds, or auditory "names" to visual objects, and to get the perceptron to perform such selective responses as are designated by the command "Name the object on the left," or "Name the color of this stimulus." The question may well be raised at this point of where the perceptron's capabilities actually stop. We have seen that the system described is sufficient for pattern recognition, associative learning, and such cognitive sets as are necessary for selective attention and selective recall. The system appears to be potentially capable of temporal pattern recognition, as well as spatial recognition, involving any sensory modality or combination of modalities. It can be shown that with proper reinforcement it will be capable of trial-and-error learning, and can learn to emit ordered sequences of responses, provided its own responses are fed back through sensory channels. Does this mean that the perceptron is capable, without further modification in principle, of such higher order functions as are involved in human speech, communication, and thinking? Actually, the limit of the perceptron's capabilities seems to lie in the area of relative judgment, and the abstraction of relationships. In its "symbolic behavior," the perceptron shows some striking similarities to Goldstein's brain-damaged patients (5). Responses to definite, concrete stimuli can be learned, even when the proper response calls for the recognition of a number of simultaneous qualifying conditions (such as naming the color if the stimulus is on the left, the shape if it is on the right). As soon as the response calls for the recognition of a relationship between stimuli (such as "Name the object left of the square." or "Indicate the pattern that appeared before the circle."), however, the THE PERCEPTRON 405
问题通常对感知机变得极其困难。仅凭统计可分性不足以提供高阶抽象的基础。此时似乎需要某种在原理上比感知机更先进的系统。结论与评价 对感知机理论研究的主要结论可总结如下:1. 在随机刺激环境中,由随机连接单元组成的系统(受上述参数约束)可以学习将特定响应与特定刺激关联起来。即使每个响应关联许多刺激,这些刺激仍能以优于随机的概率被识别,尽管它们可能彼此相似并激活系统的许多相同感觉输入。2. 在这种“理想环境”中,正确响应的概率随着学习刺激数量的增加而降低到其原始随机水平。3. 在这种环境中,不存在泛化的基础。4. 在“分化环境”中,每个响应与一类互相关或“相似”刺激相关联,系统所学特定刺激关联被正确保持的概率通常随着系统学习刺激数量的增加而趋近于一个优于随机的渐近线。通过增加系统中的关联单元数量,该渐近线可以任意接近 1。5. 在分化环境中,从未见过的刺激被正确识别并关联到其适当类别的概率(正确泛化的概率)趋近于与先前强化刺激的正确响应概率相同的渐近线。如果满足不等式\(P_{ci}^2 < P_a < P_{en}\)(针对所考虑刺激类别),则该渐近线优于随机。6. 系统的性能可以通过使用轮廓敏感的投影区域和使用二进制响应系统来改善,其中每个响应或“比特”对应于刺激的某个独立特征或属性。7. 在二值强化系统中可以进行试错学习。8. 刺激模式和响应的时间组织可以由仅使用原始统计可分性原理扩展的系统学习,而无需引入系统组织的任何重大复杂性。9. 感知机的记忆是分布式的,即任何关联都可能使用系统中的大部分单元,并且移除部分关联系统不会对任何单个辨别或关联的性能产生明显影响,但会表现为所有已学关联的普遍缺陷。10. 简单的认知集、选择性回忆以及对给定环境中存在的类别的自发识别是可能的。然而,空间和时间中关系的识别似乎代表了感知机形成认知抽象能力的极限。心理学家,特别是学习理论家现在可能会问:“除了 Hull、Bush 和 Mosteller 等人的定量理论,或 Hebb 等人的生理学理论已经完成的工作之外,当前的理论取得了什么成就?”当然,当前的理论仍然过于原始,不能被视为人类学习现有理论的成熟对手。不过,作为第一近似,其主要成就可以表述如下:对于给定的组织模式(\(alpha\)、\(beta\)或\(gamma\);S 或 n;单值或二值),学习、知觉辨别和泛化的基本现象可以完全由六个基本物理参数预测,即:
problem generally becomes excessively difficult for the perceptron. Statistical separability alone does not provide a sufficient basis for higher order abstraction. Some system, more advanced in principle than the perceptron, seems to be required at this point. CONCLUSIONS AND EVALUATION The main conclusions of the theoretical study of the perceptron can be summarized as follows: 1. In an environment of random stimuli, a system consisting of randomly connected units, subject to the parametric constraints discussed above, can learn to associate specific responses to specific stimuli. Even if many stimuli are associated to each response, they can still be recognized with a better-than-chance probability, although they may resemble one another closely and may activate many of the same sensory inputs to the system. 2. In such an "ideal environment," the probability of a correct response diminishes towards its original random level as the number of stimuli learned increases. 3. In such an environment, no basis for generalization exists. 4. In a "differentiated environment," where each response is associated to a distinct class of mutually correlated, or "similar" stimuli, the probability that a learned association of some specific stimulus will be correctly retained typically approaches a better-than-chance asymptote as the number of stimuli learned by the system increases. This asymptote can be made arbitrarily close to unity by increasing the number of association cells in the system. 5. In the differentiated environment, the probability that a stimulus which has not been seen before will be correctly recognized and associated to its appropriate class (the probability of correct generalization) approaches the same asymptote as the probability of a correct response to a previously reinforced stimulus. This asymptote will be better than chance if the inequality \(P_{ci}^2 < P_a < P_{en}\) is met, for the stimulus classes in question. 6. The performance of the system can be improved by the use of a contour-sensitive projection area, and by the use of a binary response system, in which each response, or "bit," corresponds to some independent feature or attribute of the stimulus. 7. Trial-and-error learning is possible in bivalent reinforcement systems. 8. Temporal organizations of both stimulus patterns and responses can be learned by a system which uses only an extension of the original principles of statistical separability, without introducing any major complications in the organization of the system. 9. The memory of the perceptron is distributed, in the sense that any association may make use of a large proportion of the cells in the system, and the removal of a portion of the association system would not have an appreciable effect on the performance of any one discrimination or association, but would begin to show up as a general deficit in all learned associations. 10. Simple cognitive sets, selective recall, and spontaneous recognition of the classes present in a given environment are possible. The recognition of relationships in space and time, however, seems to represent a limit to the perceptron's ability to form cognitive abstractions. Psychologists, and learning theorists in particular, may now ask: "What has the present theory accomplished, beyond what has already been done in the quantitative theories of Hull, Bush and Mosteller, etc., or physiological theories such as Hebb's?" The present theory is still too primitive, of course, to be considered as a full-fledged rival of existing theories of human learning. Nonetheless, as a first approximation, its chief accomplishment might be stated as follows: For a given mode of organization (\(alpha\), \(beta\), or \(gamma\); S or n; monovalent or bivalent) the fundamental phenomena of learning, perceptual discrimination, and generalization can be predicted entirely from six basic physical parameters, namely:
In all of the systems analyzed up to this point, the increments of value gained by an active A-unit, as a result of reinforcement or experience, have always been positive, in the sense that an active unit has always gained in its power to activate the responses to which it is connected. In the gamma-system, it is true that some units lose value, but these are always the inactive units, the active ones gaining in proportion to their rate of activity. In a bivalent system, two types of reinforcement are possible (positive and negative), and an active unit may either gain or lose in value, depending on the momentary state of affairs in the system. If the positive and negative reinforcement can be controlled by the application of ex-ternal stimuli, they become essentially equivalent to "reward" and "punish-ment," and can be used in this sense by the experimenter. Under these conditions, a perceptron appears to be capable of trial-and-error learning. Abivalent system need not necessarily involve the application of reward and punishment, however. If a binary-coded response system is so organized that there is a single response or response-pair to represent each "bit," or stimulus characteristic that is learned, with positive feedback to its own source-set if the response is "on," and negative feedback (in the sense that active A-units will lose rather than gain in value) if the response is "off," then the system is still bivalent in its characteristics. Such a bivalent THE PERCEPTRON 403
x:每个 A 单元的兴奋性连接数;
x: the number of excitatory connections per A-unit,
y:每个 A 单元的抑制性连接数;
y: the number of inhibitory connections per A-unit,
θ:A 单元的预期阈值,ω:A 单元连接的 R 单元比例;
θ: the expected threshold of an A-unit, ω: the proportion of R-units to which an A-unit is connected,
NA:系统中的 A 单元数量;
NA: the number of A-units in the system, and
NR:系统中的 R 单元数量。
NR: the number of R-units in the system.
Ns(感觉单元的数量)如果非常小就会变得重要。假设系统开始时所有单元都处于统一的值状态;否则,还需要初始值分布。
Ns (the number of sensory units) becomes important if it is very small. It is assumed that the system begins with all units in a uniform state of value; otherwise the initial value distribution would also be required.
上述每个参数都是一个明确定义的物理变量,本身是可测量的,独立于我们试图预测的行为和感知现象。
Each of the above parameters is a clearly defined physical variable, which is measurable in its own right, independently of the behavioral and perceptual phenomena which we are trying to predict.
直接由于其在物理变量上的基础,当前系统在三个主要方面远远超越了现有的学习和行为理论:简练性、可验证性、以及解释力和普适性。让我们依次考虑这些点。1. 简练性。本质上,该系统使用的所有基本变量和定律都已经存在于物理和生物科学的结构中,因此我们只需要假设一个假设变量(或构念),我们称之为 V,即关联细胞的“值”;这是一个必须符合某些明确可述的功能特征的变量,并且假定它有一个潜在可测量的物理相关物。2. 可验证性。以前的定量学习理论,显然无一例外,都有一个重要的共同特征:它们都基于在特定情境下对行为的测量,使用这些测量(经过理论处理)来预测
As a direct consequence of its foundation on physical variables, the present system goes far beyond existing learning and behavior theories in three main points: parsimony, verifiability, and explanatory power and generality. Let us consider each of these points in turn. 1. Parsimony. Essentially all of the basic variables and laws used in this system are already present in the structure of physical and biological science, so that we have found it necessary to postulate only one hypothetical variable (or construct) which we have called V, the "value" of an association cell; this is a variable which must conform to certain functional characteristics which can clearly be stated, and which is assumed to have a potentially measurable physical correlate. 2. Verifiability. Previous quantitative learning theories, apparently without exception, have had one important characteristic in common: they have all been based on measurements of behavior, in specified situations, using these measurements (after theoretical manipulation) to predict
在其他情境下的行为。这样的程序,归根结底,是一种曲线拟合和外推的过程,希望描述一组曲线的常数在其他情境下对其他曲线也适用。虽然这种外推在严格意义上不一定是循环论证,但它具有循环逻辑的许多困难,特别是当用作行为的“解释”时。这种外推在新的情境中难以证明其合理性,并且已经表明,如果基本常数和参数需要针对任何经验上失效的情境(例如从白鼠到人类的变化)重新推导,那么基本的“理论”本质上是不可反驳的,就像任何成功的曲线拟合方程不可反驳一样。事实上,心理学家普遍承认,试图“反驳”目前使用的任何主要学习理论几乎没有意义,因为通过扩展或参数变化,它们 THE PERCEPTRON 407
behavior in other situations. Such a procedure, in the last analysis, amounts to a process of curve fitting and extrapolation, in the hope that the constants which describe one set of curves will hold good for other curves in other situations. While such extrapolation is not necessarily circular, in the strict sense, it shares many of the logical difficulties of circularity, particularly when used as an "explanation" of behavior. Such extrapolation is difficult to justify in a new situation, and it has been shown that if the basic constants and parameters are to be derived anew for any situation in which they break down empirically (such as change from white rats to humans), then the basic "theory" is essentially irrefutable, just as any successful curve-fitting equation is irrefutable. It has, in fact, been widely conceded by psychologists that there is little point in trying to "disprove" any of the major learning theories in use today, since by extension, or a change in parameters, they THE PERCEPTRON 407
都证明能够适应任何特定的经验数据。这体现在越来越普遍的态度中,即理论模型的选择主要是个人审美偏好或偏见的问题,每个科学家都有权拥有自己喜欢的模型。在考虑这种方法时,人们想起了基斯季亚科夫斯基的一句话:“给我七个参数,我就能拟合一头大象。”这显然
have all proved capable of adapting to any specific empirical data. This is epitomized in the increasingly common attitude that a choice of theoretical model is mostly a matter of personal aesthetic preference or prejudice, each scientist being entitled to a favorite model of his own. In considering this approach, one is reminded of a remark attributed to Kistiakowsky, that "given seven parameters, I could fit an elephant." This is clearly
这与一个系统的情况不同,在该系统中,自变量或参数可以独立测量。
It is not the case with a system in which the independent variables, or parameters, can be measured independently.
预测行为的。在这样的系统中,如果当前使用的参数导致不当结果,则不可能“强制”拟合经验数据。在当前的理论中,在新情况下无法拟合曲线将清楚地表明,要么理论有误,要么经验测量有误。因此,如果这样的理论在重复测试中成立,那么我们可以对其有效性和普遍性更有信心,而必须针对每种情况量身定制的理论则不然。3. 解释力与普遍性。
of the predicted behavior. In such a system, it is not possible to "force" a fit to empirical data, if the parameters in current use should lead to improper results. In the current theory, a failure to fit a curve in a new situation would be a clear indication that either the theory or the empirical measurements are wrong. Consequently, if such a theory does hold up for repeated tests, we can be considerably more confident of its validity and of its generality than in the case of a theory which must be hand-tailored to meet each situation. 3. Explanatory power and generality.
当前理论源于基本物理变量,并不针对任何特定的有机体或学习情境。原则上它可以推广以涵盖任何已知物理参数系统中的任何行为形式。建立在这些基础上的学习理论应该比以往提出的任何理论都强大得多。它不仅会告诉我们任何已知有机体可能发生的行为,还会允许合成行为系统以满足特殊需求。其他学习理论在推广时往往变得越来越定性化。因此,一组描述奖励对白鼠 T 迷宫学习影响的方程,当我们试图将其推广到任何物种和任何情境时,就简化为一个陈述:受奖励的行为往往以递增的概率发生。这里提出的理论通过普遍性并未失去任何精确性。Donald Hebb(7)提出的理论试图通过展示心理功能如何从神经生理学理论推导出来,来避免基于行为模型的这些困难。在尝试实现这一目标时,Hebb 的方法哲学与我们的相当接近,他的工作一直是这里提出的大部分内容的灵感来源。然而,Hebb 从未真正实现一个能够从生理系统预测行为(或任何心理学数据)的模型。他的生理学更像是关于可能支撑行为的有机基质的一种建议,以及展示生物物理学与心理学之间桥梁合理性的尝试。当前理论代表了这种桥梁的首次实际完成。通过使用前面部分的方程,可以从神经学变量预测学习曲线,同样也可以从学习曲线预测神经学变量。这座桥梁在反复穿越中表现如何还有待观察。与此同时,这里报告的理论清楚地展示了定量统计方法在认知系统组织中的可行性和成果。通过研究像感知机这样的系统,希望那些对所有信息处理系统(包括机器和人)都共同的基本组织法则最终能被理解。
The present theory, being derived from basic physical variables, is not specific to any one organism or learning situation. It can be generalized in principle to cover any form of behavior in any system for which the physical parameters are known. A theory of learning, constructed on these foundations, should be considerably more powerful than any which has previously been proposed. It would not only tell us what behavior might occur in any known organism, but would permit the synthesis of behaving systems, to meet special requirements. Other learning theories tend to become increasingly qualitative as they are generalized. Thus a set of equations describing the effects of reward on T-maze learning in a white rat reduces simply to a statement that rewarded behavior tends to occur with increasing probability, when we attempt to generalize it from any species and any situation. The theory which has been presented here loses none of its precision through generality. The theory proposed by Donald Hebb (7) attempts to avoid these difficulties of behavior-based models by showing how psychological functioning might be derived from neuro-physiological theory. In his attempt to achieve this, Hebb's philosophy of approach seems close to our own, and his work has been a source of inspiration for much of what has been proposed here. Hebb, however, has never actually achieved a model by which behavior (or any psychological data) can be predicted from the physiological system. His physiology is more a suggestion as to the sort of organic substrate which might underlie behavior, and an attempt to show the plausibility of a bridge between biophysics and psychology. The present theory represents the first actual completion of such a bridge. Through the use of the equations in the preceding sections, it is possible to predict learning curves from neurological variables, and likewise, to predict neurological variables from learning curves. How well this bridge stands up to repeated crossings remains to be seen. In the meantime, the theory reported here clearly demonstrates the feasibility and fruitfulness of a quantitative statistical approach to the organization of cognitive systems. By the study of systems such as the perceptron, it is hoped that those fundamental laws of organization which are common to all information handling systems, machines and men included, may eventually be understood.
参考文献 1. ASHBY, W. R. 《大脑设计》。纽约:Wiley, 1952. 2. CULBERTSON, J. T. 《意识与行为》。迪比克,爱荷华:Wm. C. Brown, 1950. 3. CULBERTSON, J. T. 《一些不经济的机器人》。收录于 C. E. Shannon & J. McCarthy(编),《自动机研究》。
REFERENCES 1. ASHBY, W. R. Design for a brain. New York: Wiley, 1952. 2. CULBERTSON, J. T. Consciousness and behavior. Dubuque, Iowa: Wm. C. Brown, 1950. 3. CULBERTSON, J. T. Some uneconomical robots. In C. E. Shannon & J. McCarthy (Eds.), Automata studies.
普林斯顿:普林斯顿大学出版社,1956. 第 99-116 页。 4. ECCLES, J. C. 《思维的神经生理学基础》。牛津:Clarendon,
Princeton: Princeton University Press, 1956. Pp. 99-116. 4. ECCLES, J. C. The neurophysiological basis of mind. Oxford: Clarendon,
5. 戈尔德斯坦,K. 从精神病理学看人性。剑桥:哈佛大学出版社,1940。6. 哈耶克,F. A. 感觉秩序。芝加哥:芝加哥大学出版社,1952。7. 赫布,D. O. 行为的组织。纽约:威利,1949。8. 克利尼,S. C. 神经网和有限自动机中的事件表示。收录于 C. E. 香农和 J. 麦卡锡(编),
5. GOLDSTEIN, K. Human nature in the light of psychopathology. Cambridge: Harvard University Press, 1940. 6. HAYEK, F. A. The sensory order. Chicago: University of Chicago Press, 1952. 7. HEBB, D. O. The organization of behavior. New York: Wiley, 1949. 8. KLEENE, S. C. Representation of events in nerve nets and finite automata. In C. E. Shannon & J. McCarthy (Eds.),
自动机研究。普林斯顿:普林斯顿大学出版社,1956。第 3-41 页。9. 科勒,W. 知觉中的关系决定。收录于 L. A. 杰弗里斯(编),
Automata studies. Princeton: Princeton University Press, 1956. Pp. 3-41. 9. KOHLER, W. Relational determination in perception. In L. A. Jeffress (Ed.),
行为的大脑机制。纽约:威利,1951。第 200-243 页。10. 麦卡洛克,W. S. 为什么心灵在大脑中。收录于 L. A. 杰弗里斯(编),
Cerebral mechanisms in behavior. New York: Wiley, 1951. Pp. 200-243. 10. McCULLOCH, W. S. Why the mind is in the head. In L. A. Jeffress (Ed.),
行为的大脑机制。纽约:威利,1951。第 42-111 页。11. 麦卡洛克,W. S. 和皮茨,W. 神经活动中固有思想的逻辑演算。数学生物物理学通报,1943,5,115-133。12. 米尔纳,P. M. 细胞组装:马克二世。心理学评论,1957,64,242-
Cerebral mechanisms in behavior. New York: Wiley, 1951. Pp. 42-111. 11. MCCULLOCH, W. S., & PITTS, W. A logical calculus of the ideas immanent in nervous activity. Bull. Math. Biophysics, 1943, 5, 115-133. 12. MILNER, P. M. The cell assembly: Mark II. Psychol. Rev., 1957, 64, 242-
13. 明斯基,M. L. 有限自动机的一些通用元素。收录于 C. E. 香农和 J. 麦卡锡(编),自动机研究。普林斯顿:普林斯顿大学出版社,1956。第 117-128 页。14. 拉舍夫斯基,N. 数学生物物理学。
13. MINSKY, M. L. Some universal elements for finite automata. In C. E. Shannon & J. McCarthy (Eds.), Automata studies. Princeton: Princeton University Press, 1956. Pp. 117-128. 14. RASHEVSKY, N. Mathematical biophysics.
芝加哥:芝加哥大学出版社,1938 年。15. ROSENBLATT, F. 感知器:认知系统中统计可分性的一种理论。布法罗:康奈尔航空实验室,报告号 VG-1196-G-1,1958 年。16. UTTLEY, A. M. 条件概率机器与条件反射。载于 C. E. Shannon & J. McCarthy (编),《自动机研究》。普林斯顿:普林斯顿大学出版社,1956 年,第 253-275 页。17. VON NEUMANN, J. 自动机的一般与逻辑理论。载于 L. A. Jeffress (编),《行为中的大脑机制》。纽约:Wiley,1951 年。
Chicago: University of Chicago Press, 1938. 15. ROSENBLATT, F. The perceptron: A theory of statistical separability in cognitive systems. Buffalo: Cornell Aeronautical Laboratory, Inc. Rep. No. VG-1196-G-1, 1958. 16. UTTLEY, A. M. Conditional probability machines and conditioned reflexes. In C. E. Shannon & J. McCarthy (Eds.), Automata studies. Princeton: Princeton University Press, 1956. Pp. 253-275. 17. VON NEUMANN, J. The general and logical theory of automata. In L. A. Jeffress (Ed.), Cerebral mechanisms in behavior. New York: Wiley, 1951.
18. VON NEUMANN, J. 概率逻辑与从不可靠组件合成可靠有机体。载于 C. E. Shannon & J. McCarthy (编),
18. VON NEUMANN, J. Probabilistic logics and the synthesis of reliable organisms from unreliable components. In C. E. Shannon & J. McCarthy (Eds.),
《自动机研究》。普林斯顿:普林斯顿大学出版社,1956 年,第 43-98 页。(收稿日期:1958 年 4 月 23 日)
Automata studies. Princeton: Princeton University Press, 1956. Pp. 43-98. (Received April 23, 1958)