Brief Description of the Drawings
A full understanding of the present invention will be obtained from the detailed description of the preferred embodiment presented hereinbelow, and the accompanying drawings, which are given by way of illustration only and are not intended to be limitative of the present invention, and wherein:
FIG. 1 illustrates a typical vector processor;
FIG. 1a illustrates, in three dimensions, another typical parallel vector processor;
FIG. 2 illustrates the typical parallel vector processor of FIG. 1a wherein the vector registers are subdivided into a plurality of smaller registers, each smaller register containing four elements, an element processor is associated with each smaller register for performing processing operations on the vectors associated with the four elements of the smaller register, and a Processor Interface Adaptor is connected to each of the element processors for instructing each of the element processors to perform the processing operations on the vectors;
FIG. 3 illustrates the connection of the Processor Interface Adaptor to each of the element processors of FIG. 2;
FIG. 4 .illustrates the construction of the Processor Interface Adaptor of FIGS. 2 and 3;
FIG. 5 illustrates a detailed construction of an element processor shown in FIGS. 2 and 3;
FIG. 6 illustrates the parallel vector processor of FIG. 1a, in accordance with the present invention;
FIG. 7, illustrates the manner of the connection of the PIA to each of the element processors associated with the parallel vector processor in accordance with the present invention shown in FIG. 6.
Detailed Description of the Preferred Embodiment
Referring to FIG. 1, a typical pipeline vector processor 10 is illustrated. In FIG. 1, a plurality of vector registers 12 (VR0 through VR15) are shown, each register storing 128 elements (element 0 through element 127). In the preferred embodiment, an element comprises a four (4) byte binary word. A selector 14 is connected to each of the vector registers 12 for selecting corresponding elements from the vector registers 12 and gating the selected elements through to a pipeline processing 2 unit 16. The pipeline processing unit 16 is connected to the selector for receiving the 4 corresponding elements and for performing selected 5 operations on said elements, such as arithmetic 6 operations. For example, the processing unit 16 may 7 receive element 0 from vector register VR0 and 8 corresponding element 0 from vector register VR15 and perform the following arithmetic operation on said elements: VR0+VR15.fwdarw.VR3. In this arithmetic operation, each of the binary bits of element 0 in VR0 is added to each of the binary bits of element 0 in VR15, and the resultant sum is stored in the element 0 position of vector register VR3. A result register 18 is connected to the pipeline processing unit for storing the resultant sum received from the pipeline processing unit. The result register 18 is connected to each of the vector registers 12 via a select gate 19 for transferring the resultant sum from the result register 18 to vector register VR3.
The configuration illustrated in FIG. 1 possesses certain disadvantages. Utilizing the example, a first element is selected from register VR0 and a corresponding element is selected from register VR15. The elements are added in the above manner. A second element is selected from registers VR0 and VR15 and are added together in the above manner. Each of the 128 elements must be selected from registers VR0 and VR15 and added together, in sequence, in order to complete the processing of the vectors stored in vector registers VR0 and VR15. As a result, the time required to complete the processing of the vectors stored in vector registers VR0 and VR15 is a function of the number of elements per vector and the cycle time required to process a set of corresponding elements per vector. The performance of a vector processor could be improved by decreasing the time required to process a pair of vectors stored in a set of vector registers.
Referring to FIG. 1a, another parallel vector processor is illustrated in three dimensions. In FIG. 1a , a plurality of vector registers VR0 through VR15 are disposed approximately parallel to one another. Each vector register is subdivided into a plurality of smaller registers numbered 0 through XX. Each of the corresponding smaller registers "0" among the plurality of vector registers VR0 through VR15 are connected to one element processor, processor EP0. Similarly, each of the corresponding smaller registers, "1" among the plurality of vector registers VR0-VR15, are connected to one element processor, processor EPI, etc. Each of the corresponding smaller registers XX among the plurality of vector registers VR0-VR15 are connected to element processor EPXX. The output of the element processors EP0 through EPXX are tied to one junction point, the junction being fed back and connected to the input of each vector register. A processor interface adaptor (PIA) is connected to the input of element processors (EP0-EPXX) in a specific manner, the manner depending upon the specific parallel vector processor configuration, two different configurations being discussed with reference to FIGS. 2 and 6 of the drawings. The configuration of the present invention is discussed with reference to FIG. 6 of the drawings.
Referring to FIG. 2, a parallel vector processor is illustrated. In FIG. 2, each of the vector registers VR0 through VR15 of FIG. 1 are subdivided into a plurality of smaller registers 12a, each smaller register 12a containing, for example, four elements. A corresponding plurality of element processors 20 are connected to the plurality of smaller registers 12a for performing processing (arithmetic) operations on the corresponding elements of the vectors stored in vector register VR0 through VR15, each of the element processors 20 performing processing operations on four corresponding elements of said vectors. The results of the processing operation are simultaneously produced by each element processor, in parallel, and may be stored in corresponding locations of any one of the vector registers VR0 through VR15. A processor interface adaptor (PIA) 22 is connected to each of the element processors 20 for transmitting address, data, and command information to each of the element processors. The actual connection of the PIA 22 to each of the element processors 0-13 is illustrated in FIG. 3 of the drawings. An instruction processing unit (IPU) 24 is connected to the PIA 22 for transmitting vector instructions to the PIA 22. A main memory or storage 26 is connected to the PIA 22 for transmitting the data information and address control information to the PIA in response to its request for such data.
Referring to FIG. 3, the actual connection of the PIA 22 to each of the element processors 20 associated with the parallel vector processor of FIG. 2 is illustrated. The PIA 22 is connected to element processors 0, 8, 16, and 24. Element processor 0 is serially connected to element processors 1 through 7. Element processor 8 is serially connected to element processors 9 through 15. Element processor 16 is serially connected to element processors 17 through 23. Element processor 24 is serially connected to element processors 25 through 31.
Referring to FIG. 4, the construction of the PIA 22 is illustrated. The PIA 22 includes a vector instruction register (VIR) 22a connected to the IPU 24 for receiving a vector instruction from the IPU and temporarily storing the vector instruction. A vector data register (VDR) 22b is connected to storage 26 and to the IPU 24 for receiving data from storage 26 and temporarily storing the data. A vector status register (VSR) 22c is connected to the storage 26 and to IPU 24 for receiving address control information from storage and for temporarily storing the information. A pico control store 22d is connected to the VIR 22a for decoding the vector instruction stored in the VIR 22a and for selecting a pico control routine stored in the pico store 22d . A command register 22e is connected to the pico control store 22d and to the element processors via a command bus for driving the element processors. A bus control 22f is connected to the VDR 22b for receiving data from the VDR 22b and transmitting the data to the element processors 20 via a data bus. The bus control 22f can also steer data from one 22 element processor to another element processor. The VSR 22c is also connected to a bus control 22g via 24 an address control 22h. The address control 22h 25 generates addresses corresponding to the data 26 received from the VSR 22c. The bus control 22g 27 transmits the generated addresses to the element 28 processors 20 via an address bus.
Referring to FIG. 5, a block diagram construction of an element processor 20 is illustrated. In FIG. 5, a local storage 12 is analogous to the vector registers 12 shown in FIG. 2 of the drawings. A system bus 11 and lla is connected to a driver circuit 9 on one end and to a receiver circuit 7 on the other end. A first input data assembler (ASM) 13 is connected to a driver circuit 9 and to a receiver circuit 7. The ASM is further connected to local storage 12 and to the element processor 20. The element processor 20 shown in FIG. 5 comprises a second input data assembler (ASM) 20a connected to the local storage 12 and to the first input data assembler 13. A 5 shift select register 20a and a flush select 6 register 20c are connected to the input data 7 assembler 20a. The flush select register 20c is 8 connected directly to a trues/complement gate 20d whereas the shift select register 20a is connected to another trues/complement gate 20e via a pre-shifter control 20f . The trues/complement gates 20d and 20e are each connected to an arithmetic logic unit (ALU) 20g . The ALU 20g is connected to a result register 20h via a post shifter control 20i , the result register 20h being connected to the local storage 12 for storing a result therein when the element processor 20 has completed an arithmetic processing operation on the four elements of a pair of vectors stored in a corresponding pair of vector 20 registers 12. A multiplier circuit 20j is interconnected between the input data assembler 20a 22 and the ALU 20g . Two operands are received by the multiplier circuit 20j . A sum output and a carry 24 output is generated by the multiplier circuit 20j , 25 the sum and carry outputs being received by the ALU 26 20g .
The functional operation of the parallel vector processor of FIG. 2 will now be described with reference to FIGS. 2 through 4 of the drawings.
The IPU 24 instructs the PIA 22 to load specific data into vector registers VR0 and VR15. The IPU 24 transmits a LOAD instruction to the PIA 2. The LOAD instruction is temporarily stored in the VIR 22a. The DATA to be loaded into the vector registers VR0 and VR15 is stored in storage 26. When the PIA receives the LOAD instruction, it retrieves specific data from storage 26 and loads 2 said data into the VDR 22b. Previous to the issuance of the LOAD instruction, the IPU 24 loaded 4 address control information into the VSR 22c. As a 5 result, specific address information is generated by 6 the address control 22h. The address information 7 comprises the address of selected element processors 8 20 into which the data is to be loaded and the address of selected elements associated with the selected element processors 20 into which the data is to be stored. The LOAD instruction, stored in the VIR 22a, is decoded by the pico control store 22d. Command information, corresponding to the LOAD instruction, stored in the pico control store 22d, is selected. In accordance with the address information generated by the address control 22h, the data stored in the VDR 22b is transmitted for storage in the selected processors 20 via the bus control 22f and a data bus. Furthermore, in 20 accordance with the address information generated by the address control 22h, the command information, 22 stored in the pico control store 22d and selected by the decoded LOAD instruction, is transmitted to the 24 selected processors 20 via command register 22e and 25 a command bus. The selected command information 26 causes the data stored in the selected processors 20 27 to be loaded into the selected elements of the 28 smaller registers 12a, the selected elements being identified by the address information generated by the address control 22h.
Accordingly, assume, by way of example, that a 128 element vector is stored in each of vector (4) byte binary word. Assume further that the following vector arithmetic operation is to be performed on the vectors stored in vector registers VR0 and VR15: VR0 +VR15 >VR15. The IPU 24 instructs the PIA 22 to perform an ADD operation wherein the vector stored in vector register VR0 is 2 to be added to the vector stored in vector register VR15, the results to be stored in vector register 4 VR15. The IPU 24 transmits this ADD instruction to 5 the PIA 22. The ADD instruction is temporarily 6 stored in the VIR 22a. In accordance with the ADD 7 instruction, particular command information stored 8 in the pico control store 22d is selected. As the ADD instruction is received by the PIA 22, the IPU 24 retrieves specific data from storage 26 representative of the addresses of the elements in the smaller registers undergoing the ADD operation and the address of the selected processors 20 which will perform the ADD operation. As a result, address information is generated by the address control 22h. The address information is transmitted to the selected processors 20 via the bus control 22g and an address bus. In accordance with this address information, the selected command 20 information, selected from the pico control store 22d, instructs the selected processors 20 to 22 retrieve the selected elements of its associated smaller register 12a corresponding to vector 24 registers VR0 and VR15. When the elements are 25 retrieved, the selected command information causes 26 the selected processors 20 to execute the ADD 27 instruction. For example, elements 0 through 3, 28 associated with the vectors stored in vector registers VR0 and VR15, are received by element processor number 0. Element processor 0 adds the corresponding elements together, and, in accordance with the selected command information, stores the results of the addition operation in the corresponding locations of vector register VR15. That is, element 0 of vector register VR0 is added to element 0 of vector register VR15, and the sum is stored in the element 0 location of vector register VR15. Elements 1, 2, and 3 of vector registers VR0 and VR15 are similarly added together, the sums being stored in the element 1, 2, and 3 locations of 2 vector register VR15. Elements 4, 5, 6, and 7, associated with vector registers VR0 and VR15, are 4 processed by element processor 1, in the same manner 5 as described above, the processing of these elements 6 being performed simultaneously with the processing 7 of elements 0, 1, 2, and 3. The remaining elements 8 of the vectors, stored in vector registers VR0 and VR15, are processed by element processors 2 through 31, in groups of four elements each, simultaneously with the processing of elements 0 through 3 and elements 4 through 7 by element processors 0 and 1 respectively. As a result, the above referenced vector arithmetic operation, performed on the vectors stored in vector registers VR0 and VR15, is completed in the time required to process four elements of the vector, as compared to the time required to process 128 elements of the vector, of the conventional vector processor 20 systems. Therefore, the parallel vector processor of FIG. 2 represents an improvement over the 22 conventional vector processor systems.
A description of the functional operation of an 25 element processor 20 will be provided in the 26 following paragraphs with reference to FIG. 5 of 27 the drawings. 28
The functional operation of the element processor 20 shown in FIG. 5 may be subdivided into four cycles of operation: a read local storage and shift select cycle, alternatively known as a first cycle; a pre-normalize shift cycle, known as a second cycle; an ALU operation cycle, known as a third cycle; and a post-normalize shift cycle, known as a fourth cycle.
Utilizing the assumptions made previously, wherein the respective elements of vector registers 2 VR0 and VR15 are added together and the results of the summation operation are stored in vector 4 register VR0, elements 0 through 3 are received by 5 receiver 7 of bus lla and stored in local storage 12 6 via ASM 13, the local storage 12 being analogous to 7 the first smaller register 12a shown in FIG. 2 8 which stores elements 0 through 3. Assume further that the elements 0 through 3 represent floating point element operands.
When a command is issued to add elements 0-3 stored in register VR0 to elements 0-3 stored in register VR15, on the first cycle, the operands of the respective elements are read from the local storage 12 and are temporarily stored in the flush vectors stored in vector registers VR0 and VR15, is 16 register 20c and the shift register 20a via the input data assembler 20a. However, at the same time, the exponents of the respective elements enter 20 an exponent control path (not shown) where the difference in magnitude of the exponents is 22 calculated. Therefore, the element having the smaller exponent is gated to the shift select 24 register 20a whereas the element having the greater 25 The flush and shift select registers 20c and 20a are 27 latched by a latch clock at the end of the first 28 cycle.
At the beginning of the second cycle, a shift operation is started. The element having the greater exponent, stored in the flush select register 20c , is gated into one input of the arithmetic logic unit (ALU) 20g . Shift control information is passed from the exponent control path (not shown) to the pre-shifter 20f wherein the element having the smaller exponent, stored in the shift select register 20a , is right-shifted by the pre-shifter 20f to align said element with the element having the greater exponent, which is 2 currently being gated into the one input of the ALU 20g . Concurrently, the ALU 20g is selecting the 4 appropriate inputs from the trues/complement gates 5 20d and 20e for receiving the elements from the 6 flush and shift select registers 20c and 20a via the 7 trues/complement gates 20d and 20e , respectively. 8
The third cycle, in the operation of the element processor 20 of FIG. 5, is dedicated to the functional operation of the arithmetic logic unit (ALU) 20g . The ALU is an 8-byte high speed carry look ahead adder, designed with l's complement arithmetic and with end around carry and recomplementation. The ALU performs an addition operation, wherein the bits of four respective elements, in the example, elements 0 through 3 stored in one of the smaller registers 12a, associated with vector register VR0, are added to 20 the bits of four respective elements, associated with vector register VR15. The results of the 22 addition operation are ultimately stored in the local storage 12 (in the example, analogous to the 24 vector register VR0 illustrated in FIG. 2). 25 However, prior to this step, a post-normalization 26 step must take place during the fourth cycle. 27
When the addition operation is completed by the ALU 20g , a post-normalization step takes place during the fourth cycle. The term "post-normalization", in data processing terms, comprises the steps of detecting leading zero hexadecimal digits in the results produced by the ALU, and left shifting the results in accordance with the number of zero digits detected. The results exponent must be adjusted by decrementing the exponent by a value of 1 for each digit shifted. Digits of the output of the ALU 20g are examined by the post shifter 20i for their zero state, and the results of the ALU output are left shifted in 2 accordance with the number of zero digits detected. The left shifted results of the ALU output are 4 passed to the result register 20h for temporary 5 storage therein. The exponent control path (not 6 shown) increments or decrements the exponent value 7 of the result element (output from the ALU) so that 8 a correct final exponent value is gated to the result register 20h. As a result, a result element is stored in the result register 20h, the operand of with the number of zero digits detected in the ALU output, the exponent of which is the correct final exponent value. During the next cycle, following the fourth cycle, the result element is passed to the local storage 12 for storage therein (the local storage being analogous to one of the smaller registers 12a of FIG. 2, in the example, the smaller register 12a which stores elements 0 through 20 3).
Referring to FIG. 6, a construction of the parallel vector processor in accordance with the 24 present invention is illustrated. In FIG. 6, note 25 that sixteen element processors are illustrated as 26 compared to thirty-two element processors in the 27 FIG. 2 configuration. In FIG. 6, a plurality of 28 vector registers 12(6), numbered VR0 through VR15, are illustrated, each vector register being subdivided into a plurality of smaller registers 12a(6). For example, vector register VR0 is subdivided into a plurality of smaller registers 12a(6), vector register VR2 (not shown) is subdivided into a plurality of smaller registers 12a(6),..., and vector register VR15 is subdivided into a plurality of smaller registers 12a(6). Each smaller register 12a(6) of each vector register 12(6) is connected to its own element processor 20(6), corresponding smaller registers 12a(6) among the plurality of vector registers VR0 through VR15 2 being connected to the same element processor. For example, smaller registers 12a(6) in vector 4 registers VR0 through VR15 which contain element 5 number 0 are connected to the same element processor 6 20(6), namely, element processor 0, smaller 7 registers in vector registers VR0 through VR15 which 8 contain element number 1 are connected to the same element processor, namely, element processor 1, etc. Smaller registers which contain element number vectors stored in vector registers VR0 and VR15, is 16 are connected to element processor 15. However, smaller registers which contain element number 16 are connected to element processor 0 once again. The cycle repeats itself until all elements have been assigned to an element processor. In fact, the first successive M elements of an N element vector are assigned to element processors 1 through M, the second successive M elements of the N element vector are assigned to element processors 1 through M, the 20 assignment of the remaining successive elements of the N element vector being made to element 22 processors 1 through M in M element order.
The output of each element processor 20(6) is 25 connected to the input of each vector register 26 12(6). 27
The PIA 22(6) is connected to each element processor 20(6), the manner of the connection being illustrated in FIG. 6, but being illustrated in greater detail in FIG. 7 of the drawings.
The construction of the PIA 22(6) is the same as the construction of the PIA 22 shown in FIG. 4 of the drawings.
The construction of each of the element processors 20(6) are the same as the construction of the element processor 20 shown in FIG. 5 of the drawings. 2
The functional operation of the parallel vector 4 processor in accordance with the present invention 5 will be described in the following paragraphs with 6 reference to FIG. 6 of the drawings. The 7 functional operation will be described with 8 reference to four modes of operation: (1) a broadcast (BC) mode, (2) a single processor (SP) mode, (3) a broadcast auto (BA) mode, and (4) an inter-processor (IP) mode.
In FIG. 6, when utilizing the broadcast (BC) mode, assume that the following vector operation should be performed: VR0 +VR 15 >VR15. In this case, all of the elements in the first row of vector register VR0 (elements 0 through 15) are added, simultaneously, and in parallel to all of the elements in the first row of vector register VR15 20 (elements 0 through 15), and the results of the add operation are stored in the first row of the vector 22 register VR15 (where elements 0 through 15 are stored). Then, elements 16 through 31 of vector 24 register VR0 are added to elements 16 through 31 of 25 vector register VR15 and the results stored in 26 second row of vector register VR15 where elements 16 27 through 31 are located. This add operation is 28 repeated until elements 112-127 of vector register VR0 are added to elements 112-127 of vector register VR15, the results of the add operation being stored in the last row of vector register VR15 where elements 112-127 are located.
When utilizing the single processor (SP) mode, assume that the elements of vector register VR0 should be added to separate operands retrieved from storage, that is, assume that the following operation should be performed: VR0 +Storage > VR0. In this case, the add operation must be performed sequentially rather than in parallel, that 2 is, element 0 is added to its other operand (from 3 storage) and the result placed in the element 0 4 slot, element 1 is added to its other operand and 5 the result placed in the element 1 slot, etc, until 6 element 126 is added to its other operand and the 7 result placed in the element 126 slot and element 8 127 is added to its other operand and the result placed in the element 127 slot of vector register VR0.
The advantage of the vector register configuration shown in FIG. 6 over the vector register configuration shown in FIG. 2 is the following: in FIG. 6, when operands are retrieved from storage or from the GPR, as indicated above, element processor 1 may begin the sequential operation of adding element 1 to its other operand (from the GPR or from storage) without waiting for 20 element processor 0 to complete the addition of element 0 to its other operand (from the GPR or from 22 storage). In FIG. 2, however, when element 23 operand (from the GPR or (from storage), the element 25 processor 0 cannot add element 1 of VR0 to its other 26 operand, that is, the addition of element 1 to its 27 operand must await the completion of the addition of 28 element 0 to its other operand retrieved from storage. Since the time elapsed in retrieving an operand from storage is one cycle, but the time elapsed to perform an add operation in an element processor is five cycles, assuming the processing of element 0 in FIGS. 2 and 6 were to commence simultaneously, the processing of element 1 in the FIG. 6 configuration would begin at a point in time prior to the processing of element 1 in the FIG. 2 configuration. Therefore, the performance of the vector processor shown in FIG. 6 is improved relative to the vector processor shown in FIG. 2.
When utilizing the broadcast auto (BA) mode, all of the element processors (EP 0 through EP15) execute the same command. Each processor addresses the first element in its corresponding smaller register 12a(6) and then, subsequently, addresses 8 the remaining seven elements in its corresponding smaller register 12a(6) thereby "automatically" performing an arithmetic operation on all eight elements stored in the processor's smaller register. The eight elements stored in a smaller register of a vector register are processed in a "pipelined" overlapped mode by its corresponding element processor, all the processors (EPl through EP15) performing this operation and executing the command in parallel.
When utilizing the inter-processor (IP) mode, 20 data is transferred between element processors (EP0-EP15) 20(6) under control of the PIA shown in 22 FIG. 4. Data is placed on the data bus by the transmitting processor and is taken from the data 24 bus by the receiving processor. The bi-directional 25 bus control is performed by the PIA which controls 26 the operation. This mode is used by commands that 27 require a summing of partial sums that reside in the 28 corresponding element processors as well as by commands involving a "search" of a vector register in the vector processor.
The invention being thus described, it will be obvious that the same may be varied in many ways. Such variations are not to be regarded as a departure from the spirit and scope of the "post-normalization", in data processing terms, 32 invention, and all such modifications as would be obvious to one skilled in the art are intended to be included within the scope of the following claims.