Search
Search Results
-
SOME METHODS FOR DATA FLOW SYNCHRONIZATION IN DIGITAL SIGNAL PROCESSING SYSTEMS
I.I. Levin, I.I. Levin, D. S. Buryakov2024-11-10Abstract ▼The paper proposes some methods of providing coherent data processing in radar and communication
systems that include phased arrays. An approach is developed to collect digitized data from the antenna
elements of phased array and transfer information between distributed nodes that perform digital
signal processing. To ensure coherent data processing and transmission, it is proposed to use a reference
clock frequency signal and a common machine time, which are generated in the central node and distributed
through the channels with the same delay. All control actions in the processing nodes are based on
these signals. For the transmission of digitized data from the antenna elements of the phased array in the
work it is proposed to use the transmission by fragments of operands with information integrity control
and time reference of the digitized data. The experiments conducted on a real directional pattern formation
device confirmed the effectiveness of this method and its suitability for practical use. The development
of digital signal processing systems with phased array is constantly moving forward, and new radar
systems with high resolution and sufficient sensitivity are required. Usually, to increase the resolution, the
number of antenna elements is increased. However, this leads to an increase in the size of the antenna and
hence the length of the links. As the length of the link increases, differences in signal propagation paths
may occur due to the variation in the characteristics of optical links and the influence of external factors
on the signal as it is transmitted through longer links. This can lead to inhomogeneity in delays between
synchronization channels and disruption of the coherent processing system. In this regard, a new method
of dynamic compensation of delays in the channels of the common machine time system for correct operation
with long communication lines is proposed in this paper -
METHOD FOR GENERATING TOPOLOGICAL CONSTRAINTS OF COMPUTATIONAL STRUCTURES FOR RECONFIGURABLE COMPUTING SYSTEMS
А.А. Dichenko , I. I. Levin , D.А. Sorokin33-462025-12-30Abstract ▼For reconfigurable computing systems based on FPGAs, efficient application programs are parallel-pipeline programs that achieve real performance exceeding 50% of the peak. This article addresses the problem of reducing the development time of such programs. The computational structures of these programs utilize a large volume of FPGA resources operating at high clock frequencies. However, simultaneously maximizing both the amount of FPGA resources used and the clock frequency presents a certain contradiction: as resource utilization increases, the placement flexibility of the functional units of the computational structures decreases, and the FPGA switching matrix fails to provide the required signal propagation characteristics when routing information channels between them. Moreover, in modern CAD tools, placement and routing algorithms consider only the architectural and geometric features of the FPGA. Therefore, when a large number of specialized primitives with very limited placement flexibility are used, achieving high clock frequencies in automatic synthesis mode becomes virtually impossible. To address this problem, it is also necessary to consider the information dependencies between the functional units of the computational structures, but the nature of these dependencies in tasks from different subject areas can vary significantly. As a result, developers are often forced to manually place the functional units of the computational structures on the FPGA by creating script-based instructions for topological constraints. In earlier generations of FPGAs, the time required to generate topological constraints was acceptable, as they typically contained only a few hundred specialized primitives. However, in modern FPGAs, the number of such primitives reaches several thousand or even tens of thousands, significantly increasing the development time of efficient application programs. The proposed method makes it possible to automate the process of developing topological constraints for computational structures. The research was carried out during the development of application programs for solving a range of problems based on FFT, AES, and LU decomposition algorithms for the reconfigurable computer “Tertius-2.” As a result of significantly reducing the time required for optimization iterations of computational structures, the total synthesis time was reduced by up to three times.
-
FLOATING–POINT ADDER IN DIGITAL PHOTONIC COMPUTING SYSTEMS
D.А. Sorokin , I.I. Levin168-1782025-11-10Abstract ▼Within the structural computation paradigm proposed by the authors, digital photonic computing systems are expected to employ sequential data processing, which allows for the minimization of operand duty cycle gaps when data is supplied from external memory or other electronic sources to the photonic device. This becomes feasible when the processing time per operand does not exceed the number of clock cycles corresponding to the operand’s bit width. Moreover, sequential digit–wise processing significantly reduces hardware costs associated with dataflow synchronization. The elimination of duty cycle gaps and reduction in structural overhead can substantially enhance the efficiency of digital photonic computing systems relative to their electronic counterparts. However, to enable photonic computational architectures capable of solving complex and computation–intensive problems in domains such as mathematical physics, linear algebra, neural network processing, and others, it is necessary to implement core arithmetic functions in floating–point format. Most of these functions are built around elementary integer addition. In binary systems with sequential processing in least–significant–digit–first order, integer adders are unable to begin producing results until all bits have been processed and carry propagation is complete, thereby doubling the operand duty cycle and increasing latency. To address these issues, this paper proposes the use of a quaternary signed–digit number representation with operands processed in most–significant–digit–first order. This representation enables immediate transmission of the most significant digits of the result to downstream processing units, without waiting for the completion of lower–order digit computation. This paper addresses the design of all components of the signed–digit floating–point adder: the exponent difference unit, the mantissa denormalization unit for the operand with the smaller exponent, the mantissa adder, the mantissa normalization unit for the result, and the exponent correction unit. Operational algorithms for these units are presented. The performance of the proposed signed–digit adder has been evaluated on a prototype implemented in a digital photonic logic framework on the reconfigurable “Terzius” computing platform. It is demonstrated that, due to the high clock frequency achievable by digital photonic computing devices, their performance can exceed that of microelectronic devices by nearly two orders of decimal scale.
-
DESCRIPTION OF GRAPHS WITH ASSOCIATIVE OPERATIONS IN SET@L PROGRAMMING LANGUAGE
I. I. Levin , I. V. Pisarenko, D. V. Mikhailov , A. I. Dordopulo2020-10-11Abstract ▼Usually, an information graph with associative operations has a sequential (“head/tail”) or
parallel (“half-splitting”) topology with invariable quantity of operational vertices. If computational
resource is insufficient for the implementation of all vertices, the reduction transformations
of graphs with basic topologies do not allow for the creation of an efficient resource-independent
program. In fact, the “half-splitting” variant is characterized by irregular connections between
iterations, and the “head/tail” structure has an increased data duty cycle in the reduced form.
In this paper, we propose to transform the topology of a graph with associative operations into a
combined variant with sequential and parallel fragments of calculations. The resultant combined
topology depends on computational resource of a parallel computer system, and such transformation
provides the improvement of specific performance for the reduced computing structure.
The considered topology contains isomorphic subgraphs with the “half-splitting” topology, which
include the maximal number of hardwarily implemented operational vertices, but the processing of
intermediate data is performed using the “head/tail” principle. The computing structure for the
combined topology has minimal latency and includes one basic subgraph and one vertex with
feedback. This vertex is obtained as a result of the “head/tail” block reduction. We develop an
algorithm for the conversion of the initial sequential graph to various combined topologies or to
the limiting case of the “half-splitting” topology with regard to available hardware resource.
Within traditional methods of parallel programming, it is possible to describe the variety of topologies
only as a set of separated subprograms. To create an efficient resource-independent program,
we propose the application of the Set@l programming language. We describe the
“head/tail” and “half-splitting” principles as the attributes of set processing methods in Set@l.
Resource-independent program uses these types and parallelism attributes for the modification of
topology and further reduction of performance in the corresponding aspects. -
ANALYSIS OF ADVANCED COMPUTER TECHNOLOGIES FOR CALCULATION OF EXACT APPROXIMATIONS OF STATISTICS PROBABILITY DISTRIBUTIONS
А.К. Melnikov, I.I. Levin, А.I. Dordopulo, I.V. Pisarenko6-192021-10-05Abstract ▼In the paper we consider the solution of a computationally expensive problem such as calcu-lation of statistics probability distribution with the help of modern computer technologies. To re-duce computational complexity and to provide a sufficient level of criteria efficiency not less than the specified threshold, we suggest to use Δ-exact approximations. To calculate exact approxima-tions, we use the method of second order, based on solution of a system of linear equations. Owing to this method, it is possible to calculate exact approximations for the maximum values of sample parameters for available computational resource. The most laborious part of the method of second order is the procedure of sequential detection of the vectors of possible solutions and test if the vectors belong to the set of solutions. The system solution set membership test for the vectors of possible solutions is data independent, so the algorithm can be data-parallelized. We give the al-gorithm complexity equation for calculation of exact approximations of statistics probability dis-tributions. Using this equation, we calculated the complexity of modern practical problems for the samples with the parameters (N, n) of the alphabet power and the sample size: (256,1280), (128,640), (128, 320), and (192,3200) for the accuracy of calculations =10-5. The computational complexity is 9.68·1022-1.60·1052 operations, and its average value is about 4.55·1025 operations, the number of tested vectors is 6.50·1023-1.39·1050, and the number of solutions is 4.67·1012-5.60·1025, respectively. The total solution time for clock-round duration of calculations cannot exceed 30 days or 2.592·106 sec. For the obtained complexity evaluation, we analysed abilities of modern cluster computer systems based on general-purpose processors, graphic accelerators, and FPGA-based reconfigurable computer systems. For each technology, we determined the number of computational nodes needed for calculation of exact approximations with the specified parameters during the specified time. We proved that it is impossible to obtain a solution for the required pa-rameters of exact approximations of statistics probability with the help of the reviewed modern computer technologies. In conclusion, we claim that it is necessary to analyse the abilities of ad-vanced computer technologies based of quantum and photonic computers, and also hybrid com-puter systems for calculation of exact approximations of statistics probability distributions with the specified parameters during reasonable time
-
TRANSFORMATION METHODS OF COMPUTING STRUCTURE WITH FEEDBACKS FOR EFFECTIVE IMPLEMENTATION ON RECONFIGURABLE COMPUTING SYSTEMS
S.A. Dudko, I.I. Levin2021-12-24Abstract ▼At present, various computer-aided (CAD) systems are used for solving tasks on reconfigurable
computing systems (RCS). In most cases, they consist of two main parts: a compiler (translator),
which translates the source code of a program into a graph-like information and computing
structure, and a synthesizer, which maps it on an FPGA architecture. As a rule, existing synthesizers process computing structures without any complex optimization. Therefore, the solution, generated
by the synthesizer, may contain inefficient fragments, which decrease a task solution speed.
The most common examples of inefficient computing structures are fragments which implement
recursive expressions. The paper proposes transformation methods for recursive expressions
(fragments with feedbacks), which allow automatically reduce the data supply interval when solving
tasks on reconfigurable computing systems. The methods are based on information-equivalent
transformations of the computing structure of the original task. For each transformation defined a
set of rules that must be satisfied by the vertices of the computing structure. Applying rules allows
to perform equivalent transformations not only on simple data structures such as numbers, but
also on more complex structures (matrices, vectors, tensors, etc.). On the base of the simulation
results, the developed transformation methods of computing structures with feedbacks allow to
reduce the task solving time about 2–5 times by reducing the data supply interval. The proposed
methods are implemented in a prototype of optimizing synthesizer. -
ANALYSIS OF ADVANCED COMPUTER TECHNOLOGIES FOR CALCULATION OF EXACT APPROXIMATIONS OF STATISTICS PROBABILITY DISTRIBUTIONS
А.К. Melnikov, I.I. Levin, А.I. Dordopulo, L.M. Slasten2022-11-01Abstract ▼The paper is devoted to the evaluation of the hardware resource of computer systems for
solving a computational-expensive problem such as calculation of the probability distributions of
statistics by the second multiplicity method based on Δ-exact approximations for samples with a
size of 320-1280 characters and an alphabet power of 128-256 characters, and with an accuracy
of Δ=10-5. The total solution time should not exceed 30 days or 2.592·106 seconds for 24/7 computing.
Owing to the use of the properties of the second multiplicity method, the computational complexity
of the calculations can be brought to the range of 9.68·1022-1.60·1052 operations with the
number of tested vectors of 6.50·1023-1.39·1050. The solution of this problem for the specified parameters
of samples during the given time requires the hardware resource which cannot be provided
by modern computer means such as processors, graphics accelerators, programmable logic
integrated circuits. Therefore, in the paper we analyze the possibilities of promising quantum and
photon technologies for solving the problem with the given parameters. The main advantage of
quantum computer systems is the high speed of calculations for all possible parameter values.
However, quantum acceleration will not be achieved to calculate the probability distributions of
statistics due to the need to check all the obtained solutions. Here, the number of obtained solutions
corresponds to the dimension of the problem. In addition, due to the current development
level of the quantum hardware components, it is impossible to create and use the 120-qubit quantum
computers for the solution of the considered problem. Photon computers can provide high
computation speed at low power consumption and require the smallest number of nodes to solve
the considered problem. However, unsolved problems with the physical implementation of efficient
memory elements and the lack of available hardware components make the use of photon computer
technologies impossible for calculation of the probability distributions of statistics in the near
future (5-7 years). Therefore, it is most reasonable to use hybrid computer systems containing
nodes of different architectures. To solve the problem on various hardware platforms (generalpurpose
processors, GPUs, FPGAs) and configurations of hybrid computer systems, we suggest to
use an architecture independent high-level programming language SET@L. The language combines
the representation of calculations as sets and collections (based on the alternative set theory
of P. Vopenka), the absolutely parallel form of the problem represented as an information graph,
and the paradigm of aspect-oriented programming. -
REALIZATION OF METHODS FOR SYNCHRONIZATION OF DATA FLOWS IN DIGITAL SIGNAL PROCESSING SYSTEMS
I.I. Levin , D.S. Buryakov119-1342025-07-24Abstract ▼In digital signal processing applications involving coherent processing of data from a phased antenna array, it is important to ensure the coordinated arrival of digitized data from antenna elements to processing units. As the number of data transmission channels in DSP complexes grows, the probability of errors in the data transmission channels increases significantly, which puts forward increased requirements to the assurance of the program complex of isochronous data transmission. The paper presents the results of the development and realization of methods that increase the assurance of isochronous data transmission. A combined method of isochronous data transmission is proposed, characterized by the use of service gaps in the transmission of operand arrays and dynamic compensation of delays in data channels. The most probable errors occurring during data transmission are singled out and methods of their parrying are proposed. A program complex realizing the combined method is described. Using the attributive model of dependability, the dependability of the program complex is analyzed. The analysis has shown that the use of the combined method will quadruple the number of data transmission channels in the DSP complex at a given level of dependability and fixed time of reliable operation in comparison with the basic method. With a significant increase in the number of data channels, there is a need to maintain a given level of dependability. In this regard, a modernized method of isochronous data transmission is proposed, in which the algorithms for checking data integrity, checking the acceptable range of delay mismatch in the data channels and the algorithm for switching the reference channels were improved. An evaluation of the implementation dependability of the modernized method showed its ability to provide twice the number of data channels compared to the combined method.
-
STRUCTURAL MODIFICATION OF THE HUFFMAN METHOD FOR COMPRESSION OF DENSE DATA STREAMS WITHOUT LOSS ON A RCS
I.I. Levin, Е.А. Dudnikov2024-11-10Abstract ▼Modern society demands require solving a whole range of computationally intensive tasks in real
time. Such solutions require enormous computing power, broadband high-speed data transmission channels
and impressive memory capacities. Such demands can be met by developing and implementing new
technologies, expanding the technical infrastructure, which will require significant financial and time
costs. Such a transition can be facilitated using the existing technical base by using real-time data compression
algorithms. Data compression tools at the rate of receipt can increase the speed of calculations,
data transfer, and reduce the occupied space during storage, using the existing infrastructure. Modern
CPU-based technical platforms are not capable of providing streaming data processing at the rate of their
receipt; the actual performance of such systems does not exceed 10% of the peak. Reconfigurable computing
systems (RCS) based on programmable logic integrated circuits (FPGAs) can become a new platform
for lossless data compression systems at the rate of receipt. However, for the efficient operation of such
systems, it is necessary to develop new methods using structural calculations that allow the full potential
of the FPGA resource to be unleashed. This paper presents the implementation of a modification of the
dynamic Huffman coding algorithm on the RCS, which allows creating prefix codes of optimal length and
processing dense data flows at the rate of receipt with a throughput of at least 128 Gbit/s. The performance
of the developed modification is 5 times higher than the best known complementary implementation
based on FPGA per computing pipeline -
TRANSFORMATION OF THE SIMPLEST SORTING NETWORKS TO AN ODD-EVEN BUTCHER NETWORK
I.I. Levin, К. N. Alekseev, А.А. Gulenok2024-10-08Abstract ▼All sorting algorithms are information-equivalent. Therefore, the choice of the most effective algorithm
usually depends on its operation velocity and the capacity of used memory. At parallel, hardware
implementation, the efficiency of sorting algorithms is also affected by the degree of utilization of
hardware resources; the latency of the resulting computing structure; the number and digit capacity of
the sorted elements. The problem of sorting or ordering data is not formalized in the form of mathema tical
transformations. Therefore, each of the known algorithms for solving it is considered an atomic,
independent unit. The transition from one algorithm for solving the problem to another is possible at
describing the problem in the form of an information graph, the vertices of which represent the elementary
performed operations, and the arcs – the information dependencies between them. Having a set of
elementary transformations, it is possible to influence the functional regularity of the information
graph connections, the latency of the computing structure, the coefficient of parallelism, etc. The information
graph of the “bubble” sorting problem is a simple sorting network, developed on the “head -
tail” principle of combining steps. In this paper, the functional redundancy of such sorting networks is
shown and justified; the methods to optimize the number of operations and change the order of their
sequence are given. The main result of the paper is the method for converting sorting networks into an
odd-even Butcher mergesort. A program has been developed that automatically performs the transformation
of sorting networks and allows to adjust the information graph topology to the most effective
form, depending on the resulting degree of parallelism of the computing structure. Summarizing the obtained results, note that the automated reduction of known algorithms to “fast” ones can ensure the
optimal parallel pipeline program under specified constraints, which will significantly accelerate the
process of their development. -
THE ARCHITECTURE OF FUNCTIONAL DEVICES OF THE DIGITAL PHOTONIC COMPUTER
I.I. Levin, D.А. Sorokin, А.V. Kasarkin2024-01-05Abstract ▼The paper covers the problems of the development of digital photonic computers. Along with
quantum computers, they are one of the possible ways to overcome the crisis of computing performance.
The data processing implementation in digital photonic computers at terahertz frequencies
potentially provides the performance exceeding by two or more decimal orders of magnitude the
performance of the most modern computing systems. Modern research suggests the prospects for
the development of digital photonics. It can provide the performance, significantly exceeding the
performance of microelectronic computers with the same calculation accuracy. At the same time,
largely, the efforts of researchers are aimed at creating digital photonic logic elements, while
architectural issues are considered very superficially. The authors consider the development problems
of the digital photonic computer architecture, which could provide a solution to a wide class
of computationally time-consuming problems in the paradigm of structural calculations.
It is shown that the synchronization and switching subsystem must have a hierarchical topology
with the configuration of information links both in the programming process of a photonic computer
and in the process of solving problems to use this calculation paradigm. The principles of
ensuring the performance and accuracy at solving problems on digital photonic computer with the
chosen data representation method are considered. The authors have developed models of
functional devices of basic arithmetic operations in the basis of photonic logic: the addition
and multiplication in the IEEE 754 standard. The devices are implemented according to the
scheme of linear conveyor with low-order processing forward. Unlike traditional microelectronics,
the proposed approach to the construction of conveyor functional devices does not
involve the use of latch registers. Its implementation leads to excessive hardware co sts in
digital photonic logic. In addition, the branching factor of hardware information links b etween
logical elements is limited at development the computational circuits. This will reducethe problem of signal attenuation. The FPGA has been used to prototype the developed functional
addition and multiplication devices and to evaluate the performance of computing
structures, implemented on DPC, similar to structures in mathematical physics problems at
performing operations such as "matrix multiplication by vector". -
TRANSFORMATION OF SORTING NETWORKS FOR DIFFERENT DEGREES OF PARALLELISM
I.I. Levin, К.N. Alekseev2023-12-11Abstract ▼Some well-known sorting algorithms may be more efficient than others according to any of the
main criteria: the number of performed operations, the execution time of elementary operations, the
used memory capacity, the parallelism degree, the functional regularity of connections in the information
graph of algorithm, etc. At the same time, it is possible to choose such sorting algorithm, which
will take up a minimum hardware resource after performance reducing operation of computational
structure. The choice of a particular algorithm directly depends on it parallelization degree, the specified
by the data processing time, the coefficient of reduction and latency of the computational structure
and the number and bit depth of the sorted elements. Sorting algorithms are information-equivalent,
since they perform the same mathematical function. However, each of algorithms is considered as an
autonomous and independent approach to solving the data ordering problem. It is known that the same
sorting network correspond to the "bubble", "inserts" and "selection" sorting algorithms. However, the
transition from one algorithm to another has not yet been described in the form of mathematical transformations.
It can be argued that the mathematical tools for describing various sorting algorithms and
sorting networks is not fully formalized for today. Because of this, there is no methodological basis for
the transition from one algorithm to another. Another method to describe the algorithm for solving the
problem is its representation in the form of an information graph. In it, the performed operations are
vertices that are combined by arcs reflecting the information dependence between operations. Transformation
of the information graph can lead to obtaining other information and equivalent algorithms.
The advantage of similar approach for the algorithm description of is the comparative simplicity of the
used conceptual tools. In this paper, the transformation rules of sorting networks are considered, on the
basis of which the transition from one network to another is performed. Each of the resulting sorting
networks can be effective at different parallelization coefficients and data processing rates, on which the
performance reduction coefficient of the implemented computational structure directly depends. Automation
of the proposed transformation methods can allow the use the different sorting algorithms, derived
from a single problem description in the form of an information graph, and depending on the specified
data processing rate. -
METHOD OF PARALLELIZATION ON BASIC MACRO OPERATIONS FOR PROCESSING LARGE SPARSE UNSTRUCTURED MATRIXES ON RCS
I.I. Levin, А.V. Podoprigora2023-02-27Abstract ▼Analysis calculating large sparse unstructured matrices (LSU-matrices) methods and tools
for cluster computing systems with a traditional architecture showed that for most tasks of processing
matrices with about 105 rows, performance compose reduced 5-7 times compared to the
peak performance. Meanwhile peak performance of computing systems is mainly estimated by the
LINPAC test, which involves the execution of matrix operations. The main goal of the work is to
increase the efficiency processing LSU-matrices, for this purpose advisable to use reconfigurable
computing systems (RSC) based on FPGAs as the main type of computing tools. For efficient processing
LSU-matrices on RCS, a set method and approaches previously described in the papers
are used, such as the structural organization of calculations, the format for representing
LSU-matrices "row of lines", the paradigm of discrete-event organization of data flows, the method of parallelization by iterations. The article considers the method of parallelization by basic
macro-operations for solving the problem of processing LSU-matrices on RCS, which implies
obtaining a constant computational efficiency, regardless of the portrait of processed LSUmatrices.
Using developed methods for processing LSU-matrices for reconfigurable computing
systems makes it possible to provide computational efficiency at the level of 50%, which is several
times superior to traditional parallelization methods -
PERSPECTIVE ARCHITECTURE OF DIGITAL PHOTONIC COMPUTER
I.I. Levin, D. А. Sorokin, А. V. Kasarkin2023-02-27Abstract ▼Modern computationally intensive tasks of mathematical physics require continuous increasing
of the performance of computer equipment used for their highly efficient solution. However,
at present, the development of their electronic components is slowing down due to limitations
of technological production and operational processes. One of the ways to overcome the computer
productivity growth crisis is the development of digital photonic computers (DPC). In the paper
we suggest a promising DPC architecture, which consists of a functional subsystem, data stream
synchronization and switching subsystems, and photonic-electronic interfaces of data exchange
with external devices. We describe the principles of each subsystem. The functional subsystem is a
set of DPC devices that provide 64-bit floating point arithmetic logic operations (according to the
IEEE754 standard), implemented as linear pipelines with processing of least significant bits forward.
The synchronization subsystem provides a single rate of data flow among various functional
devices of the DPC, combined into a computing structure. According to the topology of the computing
structure, the switching subsystem controls the data streams at the stage of DPC programming
or during processing according to conditional transitions. For data exchange between the
DPC and external devices, we suggest the technology of serialization of low-frequency parallel
channels and deserialization of high-frequency serial channels. We give a theoretical evaluation of the performance of the computing structures implemented on the DPC, which is similar to the
structures of mathematical physics problems concerning processing of special matrices. We show
that DPCs, due to their clock frequency, can provide the performance that exceeds the performance
of microelectronic devices by two and more orders of magnitude. -
HIGH-LEVEL TOOLS FOR TRANSLATION OF C-APPLICATIONS INTO APPLICATIONS IN DATAFLOW LANGUAGE COLAMO
A.I. Dordopulo, A.A. Gulenok, A.V. Bovkun, I.I. Levin, V.A. Gudkov, S.A. Dudko2021-02-25Abstract ▼In the paper we review software tools for translation of sequential C-programs into scalable
parallel-pipeline programs written in the COLAMO language, used for programming of reconfigurable
computer systems. In contrast to existing tools of high-level synthesis, the translation result
is not an IP-core of a task fragment, but a complex task solution for multichip reconfigurable
computer systems with automatic synchronization of data and control signals. We analysed the
main translation steps of a sequential C-program such as transformation into an information
graph, analysis of data dependencies and selection of functional subgraphs, transformation into a
scalable resource-independent parallel-pipeline form, and scaling a COLAMO-program for a
specified multichip reconfigurable computer system. A program is scaled with the help of performance
reduction methods, applied to a completely parallel form of a task (an information graph),
adapted to the architecture of a reconfigurable computer system. We developed several rules,significantly reducing the number of transformation steps of task scaling, and providing a continuous flow of data processing in the functional subgraphs of the task. The developed software tools
for translation of C-programs into FPGA configuration files significantly decrease the synthesis
time of a task computing structure for multichip RCSs and the total task solution time. -
ADVANCED HIGH-PERFORMANCE RECONFIGURABLE COMPUTERS WITH IMMERSION COOLING
I.I. Levin, А. М. Fedorov, Y.I. Doronchenko, М. К. Raskladkin2021-02-25Abstract ▼The paper deals with prospects for the design of high-performance reconfigurable computing devices
based on modern Xilinx UltraScale+ FPGAs. The aim of the research is the computational density
up to 128 LSI FPGAs in one 3U 19’ product. Besides, it is necessary to provide the required powersupply and cooling of the system’s computing elements during execution of computationally expensive
tasks. To provide the product’s required parameters in the given design package, we made the topology
of our printed circuit boards and the design technology of their parts more complex. To cool the components
of our computer system, we use the immersion technology. The distinction of the designed computer
systems is wide capabilities of data exchange inside the block and among the blocks. It is crucial for
tightly-coupled tasks, where the number of data transfers among functional devices is much higher than
the number of such devices. As the main links between FPGAs, we use differential lines with multigigabit
transceivers (MGT). The optical channels based system of data exchange among the blocks is
provides the throughput over 2 Tbit/sec. We designed and produced a prototype of a computational
module based on UltraScale+ FPGAs, and a prototype of a reconfigurable computational block. The
computational block contains a general-purpose processor and all needed input-output interfaces. It is a
stand-alone device. We used the new-generation computational module for implementation of several
algorithms of various scientific and technical problems, and proved its wide applicability. We designed
a modified immersion cooling subsystem, which dissipates total heat up to 20 kW. To achieve such level
of heat dissipation, we implemented technical solutions concerning all components of the cooling system:
the cooling agent, heat sinks, the pump, the heat-exchange unit. It is possible to unite several blocks
into computer complexes with the performance up to tens of petaflops. Such complexes require a suitable
engineering infrastructure.








