TWI864600B - Memory management device and method applied to intelligence processing unit - Google Patents

Memory management device and method applied to intelligence processing unit Download PDF

Info

Publication number
TWI864600B
TWI864600B TW112105770A TW112105770A TWI864600B TW I864600 B TWI864600 B TW I864600B TW 112105770 A TW112105770 A TW 112105770A TW 112105770 A TW112105770 A TW 112105770A TW I864600 B TWI864600 B TW I864600B
Authority
TW
Taiwan
Prior art keywords
circuit
mapping
virtual
memory
original data
Prior art date
Application number
TW112105770A
Other languages
Chinese (zh)
Other versions
TW202435080A (en
Inventor
劉健
Original Assignee
大陸商星宸科技股份有限公司
Priority date (The priority date is an assumption and is not a legal conclusion. Google has not performed a legal analysis and makes no representation as to the accuracy of the date listed.)
Filing date
Publication date
Application filed by 大陸商星宸科技股份有限公司 filed Critical 大陸商星宸科技股份有限公司
Priority to TW112105770A priority Critical patent/TWI864600B/en
Publication of TW202435080A publication Critical patent/TW202435080A/en
Application granted granted Critical
Publication of TWI864600B publication Critical patent/TWI864600B/en

Links

Images

Landscapes

  • Memory System Of A Hierarchy Structure (AREA)
  • Bus Control (AREA)

Abstract

A memory management device includes a pre-fetch circuit, a setting circuit, and a mapping circuit. The pre-fetch circuit receives original data via a direct memory access (DMA) circuit, and the original data indicates a mapping relation between a first virtual address and physical addresses of a memory. The setting circuit analyzes the original data to sequentially map the physical addresses into second virtual addresses including the first virtual address and issues a write request. The mapping circuit stores a mapping relation between the physical addresses and the second virtual addresses to be a first mapping table and utilizes the first mapping table according to at least read request corresponding to at least one channel of the DMA circuit to access the memory.

Description

應用於智慧型處理器的記憶體管理裝置與方法Memory management device and method for intelligent processor

本案是關於記憶體管理裝置與方法,尤其是可改善智慧型處理器的記憶體管理效率之記憶體管理裝置與方法。 This case is about a memory management device and method, and in particular, a memory management device and method that can improve the memory management efficiency of an intelligent processor.

隨著人工智慧技術的發展,智慧型處理器的使用場景越來越多樣化。在現有技術中,可提高智慧處理器的內部儲存空間,以滿足智慧型處理器對於該些場景所需要的記憶體頻寬存取需求。在現有技術中,智慧型處理器的記憶體管理可能會產生碎片化的資料存取(通常涉及多個不連續的實體地址),或是需要完整搜尋緩衝區來取得記憶體的實體地址。如此,會使得記憶體管理效率不彰,從而影響指令處理效率。 With the development of artificial intelligence technology, the use scenarios of smart processors are becoming more and more diverse. In the existing technology, the internal storage space of the smart processor can be increased to meet the memory bandwidth access requirements of the smart processor for these scenarios. In the existing technology, the memory management of the smart processor may generate fragmented data access (usually involving multiple discontinuous physical addresses), or require a complete search buffer to obtain the physical address of the memory. This will make the memory management efficiency inefficient, thereby affecting the instruction processing efficiency.

於一些實施態樣中,本案的目的之一在於提供一種記憶體管理裝置與方法,以可改善先前技術的缺點。 In some implementations, one of the purposes of this case is to provide a memory management device and method to improve the shortcomings of the prior art.

於一些實施態樣中,應用於智慧處理器的記憶體管理裝置包含預取電路、設定電路以及映射電路。預取電路經由一直接記憶體存取電路取得一原始資料,該原始資料指示一第一虛擬地址與一記憶體的複數個實體地址之間的映射關係。設定電路解析該原始資料以將該些實體地址依序映射至包含該第一虛擬地址的複數個第二虛擬地址並發出一寫入請求。映射電路根據該寫入請求儲存該些實體地址與該些第二虛擬地址之間的映射關係為一第一映射表,並根據對應於該直接記憶體存取電路的至少一通道之至少一讀取請求利用該第一映射表以存取該記憶體。 In some embodiments, a memory management device for an intelligent processor includes a pre-fetch circuit, a setting circuit, and a mapping circuit. The pre-fetch circuit obtains an original data through a direct memory access circuit, and the original data indicates a mapping relationship between a first virtual address and a plurality of physical addresses of a memory. The setting circuit parses the original data to sequentially map the physical addresses to a plurality of second virtual addresses including the first virtual address and issues a write request. The mapping circuit stores the mapping relationship between the physical addresses and the second virtual addresses as a first mapping table according to the write request, and uses the first mapping table to access the memory according to at least one read request corresponding to at least one channel of the direct memory access circuit.

於一些實施態樣中,記憶體管理方法包含下列操作:經由一直接記憶體存取電路取得一原始資料,該原始資料指示一第一虛擬地址與一記憶體的複數個實體地址之間的映射關係;解析該原始資料以將該些實體地址依序映射至包含該第一虛擬地址的複數個第二虛擬地址並發出一寫入請求;以及根據該寫入請求儲存該些實體地址與該些第二虛擬地址之間的映射關係為一第一映射表,並根據對應於該直接記憶體存取電路的至少一通道之至少一讀取請求利用該第一映射表以存取該記憶體。 In some implementations, the memory management method includes the following operations: obtaining a raw data through a direct memory access circuit, the raw data indicating a mapping relationship between a first virtual address and a plurality of physical addresses of a memory; parsing the raw data to sequentially map the physical addresses to a plurality of second virtual addresses including the first virtual address and issuing a write request; and storing the mapping relationship between the physical addresses and the second virtual addresses as a first mapping table according to the write request, and using the first mapping table to access the memory according to at least one read request corresponding to at least one channel of the direct memory access circuit.

有關本案的特徵、實作與功效,茲配合圖式作較佳實施例詳細說明如下。 The features, implementation and effects of this case are described in detail below with reference to the diagrams for a preferred embodiment.

本文所使用的所有詞彙具有其通常的意涵。上述之詞彙在普遍常用之字典中之定義,在本案的內容中包含任一於此討論的詞彙之使用例子僅為示例,不應限制到本案之範圍與意涵。同樣地,本案亦不僅以於此說明書所示出的各種實施例為限。 All terms used in this article have their usual meanings. The definitions of the above terms in commonly used dictionaries and the use examples of any term discussed herein in the content of this case are only examples and should not limit the scope and meaning of this case. Similarly, this case is not limited to the various embodiments shown in this specification.

關於本文中所使用之『耦接』或『連接』,均可指二或多個元件相互直接作實體或電性接觸,或是相互間接作實體或電性接觸,亦可指二或多個元件相互操作或動作。如本文所用,用語『電路』可為由至少一個電晶體與/或至少一個主被動元件按一定方式連接以處理訊號的裝置。 As used herein, "coupling" or "connection" may refer to two or more components making physical or electrical contact directly or indirectly, or two or more components operating or acting on each other. As used herein, the term "circuit" may refer to a device that is composed of at least one transistor and/or at least one active and passive component connected in a certain manner to process signals.

圖1為根據本案一些實施例繪製一種記憶體管理裝置100的示意圖。在一些實施例中,記憶體管理裝置100可應用於一智慧型處理器(Intelligence Processing Unit)內,以管理該智慧型處理器的內部記憶體來提高內部記憶體的利用率。 FIG1 is a schematic diagram of a memory management device 100 according to some embodiments of the present invention. In some embodiments, the memory management device 100 can be applied to an intelligent processor (Intelligence Processing Unit) to manage the internal memory of the intelligent processor to improve the utilization rate of the internal memory.

記憶體管理裝置100包含預取(pre-fetch)電路110、設定電路120、映射電路130以及控制電路140。預取電路110耦接至直接記憶體存取 (Direct Memory Access)電路100A以存取一外部記憶體(例如為,但不限於,動態隨機存取記憶體)與/或智慧型處理器中的一快取記憶體。預取電路110可經由直接記憶體存取電路100A取得原始資料OD。於一些實施例中,原始資料OD可指示第一虛擬地址與一記憶體(例如為前述的外部記憶體或快取記憶體)中的多個實體地址之間的映射關係。關於原始資料OD的設置方式將於後參照圖2說明。 The memory management device 100 includes a pre-fetch circuit 110, a setting circuit 120, a mapping circuit 130, and a control circuit 140. The pre-fetch circuit 110 is coupled to a direct memory access circuit 100A to access an external memory (such as, but not limited to, a dynamic random access memory) and/or a cache memory in an intelligent processor. The pre-fetch circuit 110 can obtain the original data OD through the direct memory access circuit 100A. In some embodiments, the original data OD can indicate a mapping relationship between a first virtual address and a plurality of physical addresses in a memory (such as the aforementioned external memory or cache memory). The setting method of the original data OD will be explained later with reference to Figure 2.

在一些實施例中,在初始階段或記憶體管理裝置100初次啟動後,預取電路110可將經由系統中的主處理器(例如為中央處理器)發出的觸發訊號TR1而進行電路內部的參數(例如,但不限於,暫存器的數值)配置。在後續的操作中,該主處理器(與/或控制電路140)可依據所要執行的命令CMD發出後續的觸發訊號TR1,以控制預取電路110經由直接記憶體存取電路100A獲得對應的原始資料OD。 In some embodiments, in the initial stage or after the memory management device 100 is first started, the pre-fetch circuit 110 can configure the parameters inside the circuit (for example, but not limited to, the value of the register) via the trigger signal TR1 sent by the main processor (for example, the central processor) in the system. In subsequent operations, the main processor (and/or the control circuit 140) can send subsequent trigger signals TR1 according to the command CMD to be executed to control the pre-fetch circuit 110 to obtain the corresponding original data OD via the direct memory access circuit 100A.

在一些實施例中,預取電路110更判斷預取電路110中的剩餘資料容量是否足夠儲存原始資料OD的一部份資料的資料量,以選擇性地儲存該部分資料,直到儲存完原始資料OD。例如,預取電路110包含預取控制電路111與緩衝器電路112。預取控制電路111受控於觸發訊號TR1,並經由直接記憶體存取電路100A依序讀取到多個部分資料(其可組成原始資料OD)。預取控制電路111可判斷緩衝器電路112當前的剩餘資料容量是否大於或等於一個部分資料的資料量,以選擇性地控制緩衝器電路112儲存該部分資料。例如,若緩衝器電路112當前的剩餘資料容量大於或等於一個部分資料的資料量,預取控制電路111可控制緩衝器電路112經由直接記憶體存取電路100A接收並儲存該部分資料。依此類推,預取控制電路111可重複執行上述操作,直到完整原始資料OD 儲存(即預先取出)至緩衝器電路112。在一些實施例中,智慧處理器中的直接記憶體存取電路100A支援多步幅(stride)、多級長度與記憶體對齊(byte align)的資料搬運。因此,在上述的預先取出原始資料OD的過程中,預取控制電路111可逐步地搬運該些部分資料,其中每個部分資料可具有固定長度,例如為,但不限於,256位元。 In some embodiments, the pre-fetch circuit 110 further determines whether the remaining data capacity in the pre-fetch circuit 110 is sufficient to store a portion of the original data OD, so as to selectively store the portion of the data until the original data OD is completely stored. For example, the pre-fetch circuit 110 includes a pre-fetch control circuit 111 and a buffer circuit 112. The pre-fetch control circuit 111 is controlled by the trigger signal TR1, and sequentially reads a plurality of partial data (which may constitute the original data OD) through the direct memory access circuit 100A. The pre-fetch control circuit 111 can determine whether the current remaining data capacity of the buffer circuit 112 is greater than or equal to the data volume of a partial data, so as to selectively control the buffer circuit 112 to store the partial data. For example, if the current remaining data capacity of the buffer circuit 112 is greater than or equal to the data volume of a partial data, the pre-fetch control circuit 111 can control the buffer circuit 112 to receive and store the partial data via the direct memory access circuit 100A. Similarly, the pre-fetch control circuit 111 can repeatedly perform the above operation until the complete original data OD is stored (i.e., pre-fetched) in the buffer circuit 112. In some embodiments, the direct memory access circuit 100A in the smart processor supports data transfer with multiple strides, multiple levels of length, and memory alignment (byte align). Therefore, in the above-mentioned process of pre-fetching the original data OD, the pre-fetch control circuit 111 can gradually transfer the partial data, wherein each partial data can have a fixed length, for example, but not limited to, 256 bits.

設定電路120解析原始資料OD以將前述的多個實體地址依序映射到多個第二虛擬地址(其包含第一虛擬地址)並發出寫入請求WR。在一些實施例中,設定電路120可由解碼器與狀態機實施,以根據原始資料OD的資料格式進行解析來獲得該些實體地址與該些第二虛擬地址之間的映射關係。關於此處之操作可參考圖3說明。在獲得上述的映射關係後,設定電路120可向映射電路130發出寫入請求WR,從而將上述的映射關係儲存為一映射表。在一些實施例中,設定電路120受控於觸發訊號TR2。控制電路140可解碼源自主處理器的一命令CMD以產生該觸發訊號TR2,以控制設定電路120對原始資料OD進行解析,以更新虛擬地址與實體地址之間的映射關係。 The setting circuit 120 parses the original data OD to sequentially map the aforementioned multiple physical addresses to multiple second virtual addresses (including the first virtual address) and issues a write request WR. In some embodiments, the setting circuit 120 can be implemented by a decoder and a state machine to parse according to the data format of the original data OD to obtain the mapping relationship between the physical addresses and the second virtual addresses. The operation here can be explained with reference to Figure 3. After obtaining the above-mentioned mapping relationship, the setting circuit 120 can issue a write request WR to the mapping circuit 130, thereby storing the above-mentioned mapping relationship as a mapping table. In some embodiments, the setting circuit 120 is controlled by the trigger signal TR2. The control circuit 140 can decode a command CMD from the main processor to generate the trigger signal TR2 to control the setting circuit 120 to parse the original data OD to update the mapping relationship between the virtual address and the physical address.

映射電路130根據寫入請求WR儲存該些實體地址與該些第二虛擬地址之間的映射關係至映射電路130的第一儲存空間以作為一第一映射表,並根據來自於直接記憶體存取電路100A的一或多個(可至少為一個)讀取請求RR1~RRN利用該第一映射表來存取前述的外部記憶體或是快取記憶體,其中該些讀取請求RR1~RRN分別對應於直接記憶體存取電路100A的不同通道。例如,當主處理器的命令CMD中有一部分的指令或操作是經由智慧處理器執行時,智慧處理器可經由直接記憶體存取電路100A發出一或多個讀取請求RR1~RRN給映射電路130。映射電路130可基於此一或多個讀取請求RR1~RRN來利 用第一映射表(或是可指示其他虛擬地址與其他實體地址之間的映射關係的其他映射表)而獲得欲使用的指令與/或資料在前述的外部記憶體或是快取記憶體中的實際儲存地址(即實體地址),進而自外部記憶體或是快取記憶體取得欲使用的指令與/或資料。 The mapping circuit 130 stores the mapping relationship between the physical addresses and the second virtual addresses in the first storage space of the mapping circuit 130 as a first mapping table according to the write request WR, and uses the first mapping table to access the external memory or cache memory according to one or more (at least one) read requests RR1-RRN from the direct memory access circuit 100A, wherein the read requests RR1-RRN correspond to different channels of the direct memory access circuit 100A. For example, when a part of the instructions or operations in the command CMD of the main processor are executed by the intelligent processor, the intelligent processor can issue one or more read requests RR1-RRN to the mapping circuit 130 through the direct memory access circuit 100A. The mapping circuit 130 can utilize the first mapping table (or other mapping tables that can indicate the mapping relationship between other virtual addresses and other physical addresses) based on the one or more read requests RR1~RRN to obtain the actual storage address (i.e., physical address) of the desired instruction and/or data in the aforementioned external memory or cache memory, and then obtain the desired instruction and/or data from the external memory or cache memory.

在一些實施例中,映射電路130更暫存該些讀取請求RR1~RRN與寫入請求WR,並仲裁該些讀取請求RR1~RRN與寫入請求WR以決定該些讀取請求RR1~RRN與寫入請求WR中每一者存取映射電路130的多個儲存空間的順序。 In some embodiments, the mapping circuit 130 further temporarily stores the read requests RR1~RRN and the write requests WR, and arbitrates the read requests RR1~RRN and the write requests WR to determine the order in which each of the read requests RR1~RRN and the write requests WR accesses the multiple storage spaces of the mapping circuit 130.

詳細而言,映射電路130包含仲裁電路131以及記憶體132。仲裁電路131包含一緩衝器131A,其可暫存該些讀取請求RR1~RRN與寫入請求WR。如此,可避免仲裁電路131在進行仲裁的過程中因請求數量過多等原因出現異常而停止接收寫入請求與/或讀取請求。仲裁電路131可執行一仲裁演算法來決定寫入請求WR與該些讀取請求RR1~RRN存取記憶體132的多個儲存空間的順序。在一些實施例中,仲裁演算法可為,但不限於,循環制(Round-Robin)演算法。記憶體132包含多個儲存空間,其可分別儲存多個映射表。例如,仲裁電路131可響應該寫入請求WR而將該些實體地址與該些第二虛擬地址之間的映射關係儲存在多個儲存空間中的第一儲存空間,以作為一第一映射表。 In detail, the mapping circuit 130 includes an arbitration circuit 131 and a memory 132. The arbitration circuit 131 includes a buffer 131A, which can temporarily store the read requests RR1~RRN and the write requests WR. In this way, it can be avoided that the arbitration circuit 131 stops receiving write requests and/or read requests due to abnormalities such as too many requests during the arbitration process. The arbitration circuit 131 can execute an arbitration algorithm to determine the order in which the write request WR and the read requests RR1~RRN access the multiple storage spaces of the memory 132. In some embodiments, the arbitration algorithm can be, but is not limited to, a Round-Robin algorithm. The memory 132 includes multiple storage spaces, which can store multiple mapping tables respectively. For example, the arbitration circuit 131 can store the mapping relationship between the physical addresses and the second virtual addresses in the first storage space among the multiple storage spaces in response to the write request WR as a first mapping table.

藉由設置仲裁電路131,直接記憶體存取電路100A與設定電路120可並行地存取映射電路130中的多個映射表,進而實現高效率的記憶體管理。例如,當設定電路120欲將第一映射表儲存到記憶體132的第一儲存空間(對應於寫入請求WR)且直接記憶體存取電路100A欲讀取記憶體132中的第二 儲存空間之映射表(例如對應於讀取請求RR1)時,由於兩者所要存取的儲存空間不同,仲裁電路131可讓設定電路120與直接記憶體存取電路100A同時存取第一與第二儲存空間。換句話說,在此情形下,設定電路120寫入第一映射表到該第一儲存空間的期間與直接記憶體存取電路100A讀取第二儲存空間的期間至少有部分重疊。如此,可提高記憶體132的存取效率,從而提高記憶體管理的效率。 By setting the arbitration circuit 131, the direct memory access circuit 100A and the setting circuit 120 can access multiple mapping tables in the mapping circuit 130 in parallel, thereby realizing efficient memory management. For example, when the setting circuit 120 wants to store the first mapping table in the first storage space of the memory 132 (corresponding to the write request WR) and the direct memory access circuit 100A wants to read the mapping table of the second storage space in the memory 132 (for example, corresponding to the read request RR1), since the storage spaces to be accessed by the two are different, the arbitration circuit 131 allows the setting circuit 120 and the direct memory access circuit 100A to access the first and second storage spaces at the same time. In other words, in this case, the period during which the circuit 120 writes the first mapping table to the first storage space and the period during which the direct memory access circuit 100A reads the second storage space at least partially overlap. In this way, the access efficiency of the memory 132 can be improved, thereby improving the efficiency of memory management.

控制電路140可解碼源自主處理器的命令CMD,並判斷命令CMD中的相依性來設定多個觸發訊號TR2以及S1~SN。控制電路140可根據多個觸發訊號TR2以及S1~SN所對應的多個虛擬暫存器值(例如為圖4B的多個虛擬暫存器值EVR1~EVR4與EVR1’~EVR4’)判斷記憶體132中的多個儲存空間的操作狀態,以設定多個觸發訊號TR2以及S1~SN的輸出順序。如前所述,觸發訊號TR2可觸發設定電路120發出寫入請求WR,且剩餘的多個觸發訊號S1~SN可觸發直接記憶體存取電路100A發出該些讀取請求RR1~RRN。多個觸發訊號S1~SN分別對應於該些讀取請求RR1~RRN。例如,觸發訊號S1可觸發直接記憶體存取電路100A發出讀取請求RR1,且觸發訊號S2可觸發直接記憶體存取電路100A發出讀取請求RR2。依此類推,應可理解多個觸發訊號S1~SN與該些讀取請求RR1~RRN之間的對應關係。關於控制電路140的詳細操作將於後參照圖4A與圖4B。 The control circuit 140 can decode the command CMD from the main processor and determine the dependencies in the command CMD to set the multiple trigger signals TR2 and S1~SN. The control circuit 140 can determine the operation status of the multiple storage spaces in the memory 132 according to the multiple virtual register values corresponding to the multiple trigger signals TR2 and S1~SN (for example, the multiple virtual register values EVR1~EVR4 and EVR1'~EVR4' in FIG. 4B) to set the output sequence of the multiple trigger signals TR2 and S1~SN. As described above, the trigger signal TR2 can trigger the setting circuit 120 to issue a write request WR, and the remaining multiple trigger signals S1~SN can trigger the direct memory access circuit 100A to issue the read requests RR1~RRN. The multiple trigger signals S1~SN correspond to the read requests RR1~RRN respectively. For example, the trigger signal S1 can trigger the direct memory access circuit 100A to issue a read request RR1, and the trigger signal S2 can trigger the direct memory access circuit 100A to issue a read request RR2. By analogy, the corresponding relationship between the multiple trigger signals S1~SN and the read requests RR1~RRN should be understood. The detailed operation of the control circuit 140 will be described later with reference to FIG. 4A and FIG. 4B .

圖2為根據本案一些實施例繪製圖1中的原始資料OD的示意圖。如圖2所示,原始資料OD包含多列資料,每一列資料包含16個資訊,且每一個資訊可包含16個位元。換言之,一列資料的長度為256個位元。 FIG2 is a schematic diagram of the original data OD in FIG1 according to some embodiments of the present invention. As shown in FIG2, the original data OD includes multiple columns of data, each column of data includes 16 information, and each information may include 16 bits. In other words, the length of a column of data is 256 bits.

以第1列至第3列的資料為例,基於從右至左且從上到下的順序,原始資料OD依序包含數量資訊(標示為TL=32)、標籤資訊(標示為Tag)、數量資訊(標示為LEN=4)、第一虛擬地址(標示為VA)之資訊、多個實體地址(依序標示為PA0~PA3)之資訊、數量資訊(標示為LEN=26)、第一虛擬地址(標示為VA)之資訊、多個實體地址之資訊(依序標示為PA0~PA25)、數量資訊(標示為LEN=1)、第一虛擬地址(標示為VA)之資訊、實體地址(標示為PA0)之資訊、數量資訊(標示為LEN=1)、第一虛擬地址(標示為VA)之資訊、實體地址(標示為PA0)之資訊以及多個無效資訊(標示為TL=0)。需特別說明,在不同資訊欄位中的虛擬地址VA與實體地址PA0~PA25可代表不同的地址。多個無效資訊為零碎的無用資訊,其可用來補齊資料以實現位元對齊。 Taking the data from the 1st to 3rd rows as an example, based on the order from right to left and from top to bottom, the original data OD includes quantity information (marked as TL=32), tag information (marked as Tag), quantity information (marked as LEN=4), information of the first virtual address (marked as VA), information of multiple physical addresses (marked as PA0~PA3 in sequence), quantity information (marked as LEN=26), first virtual address (marked as The information of the first virtual address (indicated as VA), the information of multiple physical addresses (indicated as PA0~PA25 in sequence), the quantity information (indicated as LEN=1), the information of the first virtual address (indicated as VA), the information of the physical address (indicated as PA0), the quantity information (indicated as LEN=1), the information of the first virtual address (indicated as VA), the information of the physical address (indicated as PA0), and multiple invalid information (indicated as TL=0). It should be specially explained that the virtual address VA and the physical address PA0~PA25 in different information fields can represent different addresses. Multiple invalid information is fragmented useless information, which can be used to fill in the data to achieve bit alignment.

數量資訊TL=32指示本次所要設定的所有實體地址的數量為32個,例如,在第1列至第3列的資料中,與實體地址相關的資訊共有32個。標籤資訊Tag可用來讓設定電路120判斷是否有讀到正確的原始資料OD。數量資訊LEN=4指示虛擬地址VA的遞增數量(相當於第二虛擬地址的數量)。例如,虛擬地址VA(相當於第一虛擬地址或第二虛擬地址的第一者)可對應於位於其左邊的實體地址PA0,虛擬地址VA+1(相當於第二虛擬地址的第二者)可對應於左邊的實體地址PA1,依此類推,虛擬地址VA+3可對應於實體地址PA3。再者,若進一步類推,在後續的數量資訊LEN=26與虛擬地址VA可指示虛擬地址VA對應到實體地址PA0,且虛擬地址VA+25對應到實體地址PA25。據此,應可理解上述多個資訊可指示出多個實體地址與多個第二虛擬地址之間的映射關係。 The quantity information TL=32 indicates that the number of all physical addresses to be set this time is 32. For example, in the data from the 1st to the 3rd columns, there are 32 pieces of information related to the physical address. The tag information Tag can be used to allow the setting circuit 120 to determine whether the correct original data OD is read. The quantity information LEN=4 indicates the incremented quantity of the virtual address VA (equivalent to the quantity of the second virtual address). For example, the virtual address VA (equivalent to the first virtual address or the second virtual address) can correspond to the physical address PA0 located on its left, the virtual address VA+1 (equivalent to the second virtual address) can correspond to the physical address PA1 on the left, and so on. The virtual address VA+3 can correspond to the physical address PA3. Furthermore, if further analogy is used, the subsequent quantity information LEN=26 and the virtual address VA can indicate that the virtual address VA corresponds to the physical address PA0, and the virtual address VA+25 corresponds to the physical address PA25. Based on this, it should be understood that the above-mentioned multiple information can indicate the mapping relationship between multiple physical addresses and multiple second virtual addresses.

基於上述的設置方式,可將多個實體地址壓縮為對應到一個虛擬地址。例如,在第一列中,數量資訊LEN=4可將一個虛擬地址VA對應到4個實體地址PA0~PA3。在一些實施例中,原始資料OD可經由外部系統或主處理器預先生成。例如,外部系統與主處理器可基於智慧處理器所執行的神經網路所應用的相關場景預先準備好原始資料OD(即,原始資料OD是離線化產生),如此,可以壓縮資料量並讓直接記憶體存取電路100A使用連續的虛擬地址,進而產生連續的指令來繳少更換映射表的頻率。 Based on the above setting method, multiple physical addresses can be compressed to correspond to one virtual address. For example, in the first row, the quantity information LEN=4 can correspond one virtual address VA to four physical addresses PA0~PA3. In some embodiments, the original data OD can be pre-generated by an external system or a main processor. For example, the external system and the main processor can prepare the original data OD in advance based on the relevant scenarios applied by the neural network executed by the smart processor (that is, the original data OD is generated offline). In this way, the data volume can be compressed and the direct memory access circuit 100A can use continuous virtual addresses, thereby generating continuous instructions to reduce the frequency of replacing the mapping table.

圖3為根據本案一些實施例繪製圖1的設定電路120所執行的操作之流程圖。在一些實施例中,下述的多個操作可用來實現狀態機的操作。 FIG3 is a flow chart showing the operations performed by the setting circuit 120 of FIG1 according to some embodiments of the present invention. In some embodiments, the following multiple operations may be used to implement the operation of the state machine.

在操作S310,從閒置狀態經由觸發訊號(例如為觸發訊號TR2)觸發而執行下一操作。在操作S320,讀取原始資料,捨棄無效資訊(例如為圖2中的無效資訊TL=0),並確認是否有正確讀取到數量資訊(例如為數量資訊TL=32)。若有正確讀取到數量資訊,執行操作S330。在操作S330,比對標籤資訊(例如為標籤資訊Tag),以確認是否有讀取到正確的原始資料。例如,當控制電路140解碼命令CMD時,控制電路140可確認與命令CMD相關的操作或指令所要使用的映射表之標籤。控制電路140可傳輸此標籤之資訊以及觸發訊號TR2給設定電路120。設定電路120可比對此標籤與原始資料OD中的標籤資訊Tag來決定是否有取得正確的原始資料OD。若有正確取得原始資料OD,執行操作S340。若沒有正確取得原始資料,執行操作S350以回報錯誤。 In operation S310, the next operation is executed from the idle state by triggering a trigger signal (e.g., trigger signal TR2). In operation S320, the original data is read, invalid information (e.g., invalid information TL=0 in FIG. 2 ) is discarded, and it is confirmed whether the quantity information is correctly read (e.g., quantity information TL=32). If the quantity information is correctly read, operation S330 is executed. In operation S330, the tag information (e.g., tag information Tag) is compared to confirm whether the correct original data is read. For example, when the control circuit 140 decodes the command CMD, the control circuit 140 can confirm the tag of the mapping table to be used for the operation or instruction related to the command CMD. The control circuit 140 can transmit the information of the tag and the trigger signal TR2 to the setting circuit 120. The setting circuit 120 can compare the tag information Tag in the original data OD to determine whether the correct original data OD is obtained. If the original data OD is obtained correctly, the operation S340 is executed. If the original data is not obtained correctly, the operation S350 is executed to report an error.

在操作S340中,解析虛擬地址之資訊以及數量資訊(例如為數量資訊TL=32)。若解析完成,執行操作S360,解析多個實體地址之資訊,以設定多個第二虛擬地址與多個實體地址之間的映射關係。例如,如前所述,在 圖2的第1列資料中,與數量資訊LEN=4的相關的多個實體地址PA0~PA3相關的映射關係為:虛擬地址VA對應於實體地址PA0,虛擬地址VA+1對應於實體地址PA1,虛擬地址VA+2對應於實體地址PA2,且虛擬地址VA+3可對應於左邊的實體地址PA3。換言之,藉由操作S340與操作S360,設定電路120可將原始資料OD中經過壓縮的映射關係(即一個虛擬地址對應到多個實體地址)還原成多個虛擬地址與多個實體地址的映射關係。若在操作S360中,所處理的實體地址的數量不同於數量資訊(例如為TL=32)的數值,代表實體地址的數量有錯。於此狀態下,執行操作S350以回報錯誤,並在回報錯誤後清除相關資料並回到閒置狀態。在操作S370中,發出寫入請求並等待仲裁完成,並在多個第二虛擬地址與多個實體地址之間的映射關係被儲存為第一映射表後回到閒置狀態。 In operation S340, the information of the virtual address and the quantity information (for example, the quantity information TL=32) are parsed. If the parsing is completed, operation S360 is executed to parse the information of the multiple physical addresses to set the mapping relationship between the multiple second virtual addresses and the multiple physical addresses. For example, as mentioned above, in the first row of data in FIG. 2, the mapping relationship associated with the multiple physical addresses PA0~PA3 associated with the quantity information LEN=4 is: the virtual address VA corresponds to the physical address PA0, the virtual address VA+1 corresponds to the physical address PA1, the virtual address VA+2 corresponds to the physical address PA2, and the virtual address VA+3 can correspond to the physical address PA3 on the left. In other words, through operation S340 and operation S360, the setting circuit 120 can restore the compressed mapping relationship in the original data OD (i.e., one virtual address corresponds to multiple physical addresses) to a mapping relationship between multiple virtual addresses and multiple physical addresses. If in operation S360, the number of processed physical addresses is different from the value of the quantity information (e.g., TL=32), it means that the number of physical addresses is wrong. In this state, operation S350 is executed to report an error, and after reporting the error, the relevant data is cleared and the idle state is returned. In operation S370, a write request is issued and arbitration is waited for to be completed, and the mapping relationship between the plurality of second virtual addresses and the plurality of physical addresses is stored as a first mapping table before returning to an idle state.

圖4A為根據本案一些實施例繪製圖1中的控制電路140的示意圖。控制電路140包含指令解碼器141、多個任務佇列電路142[0]~142[N]、多個觸發電路143[0]~143[N]、外部虛擬暫存器(external virtual register)佇列電路144以及相依性確認電路145。 FIG4A is a schematic diagram of the control circuit 140 in FIG1 according to some embodiments of the present invention. The control circuit 140 includes an instruction decoder 141, a plurality of task queue circuits 142[0]~142[N], a plurality of trigger circuits 143[0]~143[N], an external virtual register queue circuit 144, and a dependency confirmation circuit 145.

多個任務佇列電路142[0]~142[N]中每一者可為,但不限於,一先入先出(FIFO)電路,任務佇列電路142[0]儲存設定電路120待執行的任務,且任務佇列電路142[1]~142[N]分別儲存直接記憶體存取電路100A的第1個至第N個通道待執行的任務。多個觸發電路143[0]~143[N]分別對應於多個任務佇列電路142[0]~142[N]設置。例如,觸發電路143[0]可根據任務佇列電路142[0]發出的要求產生觸發訊號TR2。觸發電路143[1]可根據任務佇列電路142[1]發出的要求產生觸發訊號S1至直接記憶體存取電路100A,以使直接記憶體存取電路100A的第1個通道發出讀取請求RR1。依此類推,觸發電路143[N]可根據任務佇 列電路142[N]發出的要求產生觸發訊號SN至直接記憶體存取電路100A,以使直接記憶體存取電路100A的第N個通道發出讀取請求RRN。 Each of the plurality of task queue circuits 142[0]-142[N] may be, but is not limited to, a first-in first-out (FIFO) circuit. The task queue circuit 142[0] stores the tasks to be executed by the setting circuit 120, and the task queue circuits 142[1]-142[N] respectively store the tasks to be executed by the 1st to Nth channels of the direct memory access circuit 100A. The plurality of trigger circuits 143[0]-143[N] are respectively configured to correspond to the plurality of task queue circuits 142[0]-142[N]. For example, the trigger circuit 143[0] may generate a trigger signal TR2 according to a request issued by the task queue circuit 142[0]. The trigger circuit 143[1] can generate a trigger signal S1 to the direct memory access circuit 100A according to the request issued by the task queue circuit 142[1], so that the first channel of the direct memory access circuit 100A issues a read request RR1. Similarly, the trigger circuit 143[N] can generate a trigger signal SN to the direct memory access circuit 100A according to the request issued by the task queue circuit 142[N], so that the Nth channel of the direct memory access circuit 100A issues a read request RRN.

指令解碼器141可解碼命令CMD以確認命令CMD所需的指令或資料,並相應地傳輸相關任務到多個任務佇列電路142[0]與142[1]~142[N]。例如,若所需的指令或資料涉及到新的映射表中所記錄的實體地址,指令解碼器141可發送任務至任務佇列電路142[0]。觸發電路143[0]可據此產生新的寫入請求WR以控制設定電路120替換現有的映射表。 The command decoder 141 can decode the command CMD to confirm the command or data required by the command CMD, and transmit the relevant tasks to multiple task queue circuits 142[0] and 142[1]~142[N] accordingly. For example, if the required command or data involves a physical address recorded in a new mapping table, the command decoder 141 can send the task to the task queue circuit 142[0]. The trigger circuit 143[0] can generate a new write request WR accordingly to control the setting circuit 120 to replace the existing mapping table.

為了確保設定電路120與控制電路140可正確地並行使用映射電路130中所儲存的多個映射表,可藉由外部虛擬暫存器(external virtual register)佇列電路144以及相依性確認電路145來判斷映射電路130中的多個儲存空間的操作狀態,以設定經由多個觸發訊號TR2與S1~SN的輸出順序。指令解碼器141可解碼命令CMD,並根據命令CMD所需的指令或資料之間的相依性設定外部虛擬暫存器佇列電路144中的多個外部虛擬暫存器值(如圖4B中的多個外部虛擬暫存器值EVR1~EVR4以及EVR1’~EVR4’),以記錄映射電路130中多個儲存空間的操作狀態(例如,是否正被設定電路120寫入中或是正由直接記憶體存取電路的一通道存取中)。相依性確認電路145可根據指令或資料之間的相依性以及多個外部虛擬暫存器值確認多個觸發訊號TR2與S1~SN的輸出順序。根據指令或資料之間的相依性,相依性確認電路145可設定多個任務佇列電路142[0]~142[N]是否可傳輸要求至多個觸發電路143[0]~143[N]。例如,相依性確認電路145可藉由中斷多個任務佇列電路142[0]~142[N]與多個觸發電路143[0]~143[N]之間的連接來設定多個觸發訊號TR2與S1~SN的輸出順序。關於此處之說明將於後參照圖4B說明。 In order to ensure that the setting circuit 120 and the control circuit 140 can correctly and concurrently use the multiple mapping tables stored in the mapping circuit 130, the operation status of the multiple storage spaces in the mapping circuit 130 can be determined by an external virtual register queue circuit 144 and a dependency confirmation circuit 145 to set the output sequence through the multiple trigger signals TR2 and S1~SN. The instruction decoder 141 can decode the command CMD, and set multiple external virtual register values (such as multiple external virtual register values EVR1~EVR4 and EVR1'~EVR4' in FIG. 4B) in the external virtual register queue circuit 144 according to the dependency between the instruction or data required by the command CMD, so as to record the operation status of multiple storage spaces in the mapping circuit 130 (for example, whether they are being written by the setting circuit 120 or being accessed by a channel of the direct memory access circuit). The dependency confirmation circuit 145 can confirm the output sequence of multiple trigger signals TR2 and S1~SN according to the dependency between the instruction or data and the multiple external virtual register values. According to the dependency between instructions or data, the dependency confirmation circuit 145 can set whether multiple task queue circuits 142[0]~142[N] can transmit requests to multiple trigger circuits 143[0]~143[N]. For example, the dependency confirmation circuit 145 can set the output sequence of multiple trigger signals TR2 and S1~SN by interrupting the connection between multiple task queue circuits 142[0]~142[N] and multiple trigger circuits 143[0]~143[N]. The description here will be described later with reference to FIG. 4B.

圖4B為根據本案一些實施例繪製圖1中的設定電路120與直接記憶體存取電路100A存取記憶體132的工作排程的示意圖。在圖4B中的例子中,相依性確認電路145可根據多個指令之間的相依性來決定設定電路120與直接記憶體存取電路100A存取記憶體132的順序。在一些實施例中,相依性確認電路145可根據智慧處理器所應用的多種相關場景事先記錄多個指令與/或資料之間的相依性,以根據當前收到的命令CMD中所包含的指令來設定多個觸發訊號TR2與S1~SN的輸出順序。 FIG4B is a schematic diagram of the work schedule of the setting circuit 120 and the direct memory access circuit 100A in FIG1 for accessing the memory 132 according to some embodiments of the present invention. In the example in FIG4B , the dependency confirmation circuit 145 can determine the order in which the setting circuit 120 and the direct memory access circuit 100A access the memory 132 according to the dependencies between the multiple instructions. In some embodiments, the dependency confirmation circuit 145 can record the dependencies between the multiple instructions and/or data in advance according to the multiple related scenarios applied by the smart processor, so as to set the output order of the multiple trigger signals TR2 and S1~SN according to the instructions included in the currently received command CMD.

舉例而言,在一些場景中,經解碼後的命令CMD包含對應於一連串的數學運算(例如可為圖像處理的運算或是卷積運算等)多個指令。例如,第1個指令可能為卷積運算,第2個指令要利用卷積運算的計算結果再進行濾波處理來產生下一個輸出。相依性確認電路145可讓直接記憶體存取電路100A先利用記憶體132的第一映射表中所指示的多個實體地址,以存取卷積運算要用到的多個指令與/或資料來進行第一層運算(例如時間t1至時間t2的操作)。接著,相依性確認電路145可讓設定電路120將記憶體132的第一映射表替換為第二映射表(例如時間t2至時間t3的操作),並控制直接記憶體存取電路100A根據該第二映射表所指示的多個實體地址,以存取濾波處理要用到的多個指令與/或資料來進行第二層運算(例如時間t3開始的操作)。 For example, in some scenarios, the decoded command CMD includes multiple instructions corresponding to a series of mathematical operations (such as image processing operations or convolution operations, etc.). For example, the first instruction may be a convolution operation, and the second instruction uses the calculation result of the convolution operation to perform filtering processing to generate the next output. The dependency confirmation circuit 145 allows the direct memory access circuit 100A to first use the multiple physical addresses indicated in the first mapping table of the memory 132 to access the multiple instructions and/or data used in the convolution operation to perform the first level operation (such as the operation from time t1 to time t2). Then, the dependency confirmation circuit 145 allows the setting circuit 120 to replace the first mapping table of the memory 132 with the second mapping table (e.g., the operation from time t2 to time t3), and controls the direct memory access circuit 100A to access the multiple instructions and/or data used in the filtering process according to the multiple physical addresses indicated by the second mapping table to perform the second-level operation (e.g., the operation starting from time t3).

詳細而言,在時間t1,直接記憶體存取電路100A的通道1正在讀取記憶體132中的第一映射表,以利用該第一映射表來獲得實體地址並自前述的外部記憶體或快取記憶體取出指令或資料來進行卷積運算。因此,直接記憶體存取電路100A的通道1所對應的外部虛擬暫存器值EVR1(其儲存於外部虛擬暫存器佇列電路144)會切換成一預設值(以斜線背景表示)以指示該記憶體132 中儲存該第一映射表的對應儲存空間處於忙碌狀態。在時間t2,直接記憶體存取電路100A的通道1結束讀取第一映射表。設定電路120可清除該對應儲存空間並將寫入第二映射表至該對應儲存空間。因此,設定電路120可將對應的外部虛擬暫存器值EVR1’與EVR2’(其儲存於外部虛擬暫存器佇列電路144)會切換成預設值,以指示該記憶體132中的該對應儲存空間處於忙碌狀態。在時間t3,設定電路120完成替換第二映射表,且直接記憶體存取電路100A的通道2正在讀取記憶體132中的第二映射表,以利用該第二映射表來獲得實體地址並自前述的外部記憶體或快取記憶體取出指令或資料來進行濾波處理。 Specifically, at time t1, channel 1 of the direct memory access circuit 100A is reading the first mapping table in the memory 132 to obtain the physical address using the first mapping table and fetch the instruction or data from the aforementioned external memory or cache memory to perform the convolution operation. Therefore, the external virtual register value EVR1 (which is stored in the external virtual register queue circuit 144) corresponding to channel 1 of the direct memory access circuit 100A is switched to a default value (indicated by a slash background) to indicate that the corresponding storage space storing the first mapping table in the memory 132 is busy. At time t2, channel 1 of the direct memory access circuit 100A finishes reading the first mapping table. The setting circuit 120 may clear the corresponding storage space and write the second mapping table to the corresponding storage space. Therefore, the setting circuit 120 may switch the corresponding external virtual register values EVR1′ and EVR2′ (which are stored in the external virtual register queue circuit 144) to default values to indicate that the corresponding storage space in the memory 132 is busy. At time t3, the setting circuit 120 completes replacing the second mapping table, and the channel 2 of the direct memory access circuit 100A is reading the second mapping table in the memory 132 to obtain the physical address using the second mapping table and fetch the instruction or data from the aforementioned external memory or cache memory for filtering processing.

直接記憶體存取電路100A的通道3與通道4存取記憶體132中第二儲存空間所儲存的第二映射表,從而獲得所需要的資料與/或資料。由於第一儲存空間不同於第二儲存空間,通道3的工作時間可與通道1的工作時間與/或設定電路120寫入第一儲存空間的時間存在部分重疊。類似地,通道2的工作時間可與通道3或4的工作時間與/或設定電路120寫入第二儲存空間的時間(即外部虛擬暫存器值EVR3’與EVR4’處於忙碌狀態的時間)存在部分重疊。 Channels 3 and 4 of the direct memory access circuit 100A access the second mapping table stored in the second storage space in the memory 132 to obtain the required data and/or data. Since the first storage space is different from the second storage space, the working time of channel 3 may partially overlap with the working time of channel 1 and/or the time when the setting circuit 120 writes to the first storage space. Similarly, the working time of channel 2 may partially overlap with the working time of channel 3 or 4 and/or the time when the setting circuit 120 writes to the second storage space (i.e., the time when the external virtual register values EVR3' and EVR4' are busy).

直接記憶體存取電路100A的通道3與通道4以及設定電路120之間的操作類似於上述操作,故不重複說明。從上述操作過程中可理解,設定電路120的替換映射表之操作不會影響通道3的處理效率,同樣的,設定電路120的替換映射表之操作不會影響通道1的處理效率。據此,藉由設置多個外部虛擬暫存器值EVR1~EVR4以及EVR1’~EVR4’,可讓設定電路120與直接記憶體存取電路100A有更高的效率並行地存取記憶體132的多個儲存空間。 The operation between channel 3 and channel 4 of the direct memory access circuit 100A and the setting circuit 120 is similar to the above operation, so it will not be repeated. From the above operation process, it can be understood that the operation of the replacement mapping table of the setting circuit 120 will not affect the processing efficiency of channel 3. Similarly, the operation of the replacement mapping table of the setting circuit 120 will not affect the processing efficiency of channel 1. Accordingly, by setting multiple external virtual register values EVR1~EVR4 and EVR1'~EVR4', the setting circuit 120 and the direct memory access circuit 100A can access multiple storage spaces of the memory 132 in parallel with higher efficiency.

圖5為根據本案一些實施例繪製記憶體管理方法500的流程圖。在操作S510,經由一直接記憶體存取電路取得一原始資料,該原始資料指示一 第一虛擬地址與一記憶體的複數個實體地址之間的映射關係。在操作S520,解析該原始資料以將該些實體地址依序映射至包含該第一虛擬地址的複數個第二虛擬地址並發出一寫入請求。在操作S530,根據該寫入請求儲存該些實體地址與該些第二虛擬地址之間的映射關係為一第一映射表,並根據對應於該直接記憶體存取電路的至少一通道之至少一讀取請求利用該第一映射表存取該記憶體。 FIG5 is a flow chart of a memory management method 500 according to some embodiments of the present invention. In operation S510, an original data is obtained through a direct memory access circuit, and the original data indicates a mapping relationship between a first virtual address and a plurality of physical addresses of a memory. In operation S520, the original data is parsed to sequentially map the physical addresses to a plurality of second virtual addresses including the first virtual address and issue a write request. In operation S530, the mapping relationship between the physical addresses and the second virtual addresses is stored as a first mapping table according to the write request, and the memory is accessed using the first mapping table according to at least one read request corresponding to at least one channel of the direct memory access circuit.

上述多個操作之說明可參照前述各個實施例,故不再重複贅述。上述記憶體管理方法500的多個操作僅為示例,並非限定需依照此示例中的順序執行。在不違背本案的各實施例的操作方式與範圍下,在記憶體管理方法500下的各種操作當可適當地增加、替換、省略或以不同順序執行(例如可以是同時執行或是部分同時執行)。 The description of the above-mentioned multiple operations can refer to the above-mentioned embodiments, so it will not be repeated. The multiple operations of the above-mentioned memory management method 500 are only examples, and are not limited to be executed in the order in this example. Without violating the operation mode and scope of each embodiment of this case, the various operations under the memory management method 500 can be appropriately added, replaced, omitted or executed in a different order (for example, they can be executed simultaneously or partially simultaneously).

綜上所述,本案一些實施例中的記憶體管理裝置與記憶體管理方法可在智慧處理器內部實現動態更新映射表、利用離線化產生具有壓縮性質的映射關係資料以及並行化的存取來提高記憶體管理的效率,從而改善智慧處理器的運行效率。 In summary, the memory management device and memory management method in some embodiments of the present invention can realize dynamic updating of mapping tables inside the smart processor, generate compressed mapping relationship data offline, and parallelize access to improve the efficiency of memory management, thereby improving the operating efficiency of the smart processor.

雖然本案之實施例如上所述,然而該些實施例並非用來限定本案,本技術領域具有通常知識者可依據本案之明示或隱含之內容對本案之技術特徵施以變化,凡此種種變化均可能屬於本案所尋求之專利保護範疇,換言之,本案之專利保護範圍須視本說明書之申請專利範圍所界定者為準。 Although the embodiments of this case are described above, these embodiments are not used to limit this case. People with ordinary knowledge in this technical field can make changes to the technical features of this case based on the explicit or implicit content of this case. All these changes may fall within the scope of patent protection sought by this case. In other words, the scope of patent protection of this case shall be subject to the scope of patent application defined in this specification.

100:記憶體管理裝置 100: Memory management device

100A:直接記憶體存取電路 100A: Direct memory access circuit

110:預取電路 110: Prefetch circuit

111:預取控制電路 111: Prefetch control circuit

112:緩衝器電路 112: Buffer circuit

120:設定電路 120: Setting circuit

130:映射電路 130: Mapping circuit

131:仲裁電路 131: Arbitration circuit

131A:緩衝器 131A: Buffer

132:記憶體 132: Memory

140:控制電路 140: Control circuit

141:指令解碼器 141: Command decoder

142[0]~142[N]:任務佇列電路 142[0]~142[N]: Task queue circuit

143[0]~143[N]:觸發電路 143[0]~143[N]: Trigger circuit

144:外部虛擬暫存器佇列電路 144: External virtual register queue circuit

145:相依性確認電路 145: Dependency confirmation circuit

500:記憶體管理方法 500:Memory management method

CMD:命令 CMD: Command

EVR1-EVR4,EVR1’-EVR4’:虛擬暫存器值 EVR1-EVR4,EVR1’-EVR4’: virtual register value

LEN=4,LEN=26,LEN=1,LEN=5,LEN=1:數量資訊 LEN=4,LEN=26,LEN=1,LEN=5,LEN=1: Quantity information

OD:原始資料 OD:Original data

PA0~PA25:實體地址 PA0~PA25: physical address

RR1~RRN:讀取請求 RR1~RRN: read request

S1~SN:觸發訊號 S1~SN: trigger signal

S310,S320,S330,S340,S350,S360,S370:操作 S310,S320,S330,S340,S350,S360,S370: Operation

S510,S520,S530:操作 S510, S520, S530: Operation

TL=32,TL=4,TL=13:數量資訊 TL=32,TL=4,TL=13: quantity information

TR1,TR2:觸發訊號 TR1,TR2: trigger signal

Tag:標籤資訊 Tag: Tag information

VA:虛擬地址 VA: Virtual Address

WR:寫入請求 WR: Write Request

t1~t3:時間 t1~t3: time

〔圖1〕為根據本案一些實施例繪製一種記憶體管理裝置的示意圖; 〔圖2〕為根據本案一些實施例繪製圖1中的原始資料的示意圖;〔圖3〕為根據本案一些實施例繪製圖1中的設定電路所執行的操作之流程圖;〔圖4A〕為根據本案一些實施例繪製圖1中的控制電路的示意圖;〔圖4B〕為根據本案一些實施例繪製圖1中的設定電路與直接記憶體存取電路存取記憶體的工作排程的示意圖;以及〔圖5〕為根據本案一些實施例繪製記憶體管理方法的流程圖。 〔Figure 1〕is a schematic diagram of a memory management device according to some embodiments of the present invention; 〔Figure 2〕is a schematic diagram of the original data in Figure 1 according to some embodiments of the present invention; 〔Figure 3〕is a flow chart of the operation performed by the setting circuit in Figure 1 according to some embodiments of the present invention; 〔Figure 4A〕is a schematic diagram of the control circuit in Figure 1 according to some embodiments of the present invention; 〔Figure 4B〕is a schematic diagram of the work schedule of the setting circuit and the direct memory access circuit in Figure 1 accessing the memory according to some embodiments of the present invention; and 〔Figure 5〕is a flow chart of the memory management method according to some embodiments of the present invention.

100:記憶體管理裝置 100: Memory management device

100A:直接記憶體存取電路 100A: Direct memory access circuit

110:預取電路 110: Prefetch circuit

111:預取控制電路 111: Prefetch control circuit

112:緩衝器電路 112: Buffer circuit

120:設定電路 120: Setting circuit

130:映射電路 130: Mapping circuit

131:仲裁電路 131: Arbitration circuit

131A:緩衝器 131A: Buffer

132:記憶體 132: Memory

140:控制電路 140: Control circuit

CMD:命令 CMD: Command

OD:原始資料 OD:Original data

RR1~RRN:讀取請求 RR1~RRN: read request

S1~SN:觸發訊號 S1~SN: trigger signal

TR1,TR2:觸發訊號 TR1,TR2: trigger signal

WR:寫入請求 WR: Write Request

Claims (10)

一種記憶體管理裝置,應用於一智慧處理器,該記憶體管理裝置包含:一預取電路,經由一直接記憶體存取電路取得一原始資料,該原始資料指示一第一虛擬地址與一記憶體的複數個實體地址之間的映射關係;一設定電路,解析該原始資料以將該些實體地址依序映射至包含該第一虛擬地址的複數個第二虛擬地址並發出一寫入請求;以及一映射電路,根據該寫入請求儲存該些實體地址與該些第二虛擬地址之間的映射關係為一第一映射表,並根據對應於該直接記憶體存取電路的至少一通道之至少一讀取請求利用該第一映射表以存取該記憶體,其中該預取電路包含:一緩衝器電路;以及一預取控制電路,判斷該緩衝器電路的一剩餘資料容量是否大於或等於該原始資料中的一部分資料之資料量,以選擇性地控制該緩衝器電路儲存該部分資料。 A memory management device is applied to an intelligent processor. The memory management device comprises: a pre-fetch circuit, which obtains an original data through a direct memory access circuit, wherein the original data indicates a mapping relationship between a first virtual address and a plurality of physical addresses of a memory; a setting circuit, which parses the original data to sequentially map the physical addresses to a plurality of second virtual addresses including the first virtual address and issues a write request; and a mapping circuit, which stores the first virtual address according to the write request. The mapping relationship between the physical addresses and the second virtual addresses is a first mapping table, and the first mapping table is used to access the memory according to at least one read request of at least one channel corresponding to the direct memory access circuit, wherein the pre-fetch circuit includes: a buffer circuit; and a pre-fetch control circuit, which determines whether a remaining data capacity of the buffer circuit is greater than or equal to the data amount of a part of the original data, so as to selectively control the buffer circuit to store the part of the data. 如請求項1之記憶體管理裝置,其中該原始資料包含一第一數量資訊、一第二數量資訊、一標籤資訊、該第一虛擬地址的資訊以及該些實體地址的資訊,其中該第一數量資訊指示該些實體地址的數量,且該第二數量資訊指示該些第二虛擬地址的數量。 A memory management device as claimed in claim 1, wherein the original data includes a first quantity information, a second quantity information, a tag information, information of the first virtual address, and information of the physical addresses, wherein the first quantity information indicates the quantity of the physical addresses, and the second quantity information indicates the quantity of the second virtual addresses. 如請求項2之記憶體管理裝置,其中該設定電路根據該第一數量資訊以及該標籤資訊判斷是否有正確預取到該原始資料,並根據該第一虛擬 地址的資訊、該第二數量資訊以及該些第二虛擬地址的資訊將該些實體地址依序映射至該些第二虛擬地址。 As in the memory management device of claim 2, the setting circuit determines whether the original data is correctly pre-fetched according to the first quantity information and the tag information, and maps the physical addresses to the second virtual addresses in sequence according to the information of the first virtual address, the second quantity information and the information of the second virtual addresses. 如請求項3之記憶體管理裝置,其中該設定電路根據該第一虛擬地址的資訊以及該第二數量資訊依序遞增該第一虛擬地址以產生該些第二虛擬地址。 A memory management device as claimed in claim 3, wherein the setting circuit sequentially increments the first virtual address according to the information of the first virtual address and the second quantity information to generate the second virtual addresses. 如請求項1之記憶體管理裝置,其中該映射電路更仲裁該寫入請求與該至少一讀取請求,以決定該寫入請求與該至少一讀取請求中每一者存取用來儲存該第一映射表的一儲存空間的順序。 A memory management device as in claim 1, wherein the mapping circuit further arbitrates the write request and the at least one read request to determine the order in which each of the write request and the at least one read request accesses a storage space used to store the first mapping table. 如請求項1之記憶體管理裝置,其中當該映射電路欲儲存該第一映射表至該映射電路中的一第一儲存空間且該至少一讀取請求欲讀取該映射電路中的一第二儲存空間時,該第一映射表儲存至該第一儲存空間的期間與該第二儲存空間響應於該至少一讀取請求被讀取的期間至少部分重疊。 A memory management device as claimed in claim 1, wherein when the mapping circuit intends to store the first mapping table to a first storage space in the mapping circuit and the at least one read request intends to read a second storage space in the mapping circuit, the period during which the first mapping table is stored in the first storage space and the period during which the second storage space is read in response to the at least one read request at least partially overlap. 如請求項1之記憶體管理裝置,其中該映射電路包含:一記憶體,包含複數個儲存空間;以及一仲裁電路,暫存該至少一讀取請求與該寫入請求,並決定該寫入請求與該至少一讀取請求存取該些儲存空間的順序,其中,該仲裁電路響應該寫入請求儲存該第一映射表至該些儲存空間中的一對應儲存空間,以替換該對應儲存空間先前儲存的一第二映射表。 A memory management device as claimed in claim 1, wherein the mapping circuit comprises: a memory comprising a plurality of storage spaces; and an arbitration circuit temporarily storing the at least one read request and the write request, and determining the order in which the write request and the at least one read request access the storage spaces, wherein the arbitration circuit stores the first mapping table to a corresponding storage space among the storage spaces in response to the write request to replace a second mapping table previously stored in the corresponding storage space. 如請求項1之記憶體管理裝置,更包含:一控制電路,解碼源自一主處理器的一指令並判斷該指令的相依性以設定複數個觸發訊號,並根據該些觸發訊號所分別對應的複數個外部虛擬暫存器值 判斷該映射電路中的複數個儲存空間的操作狀態以設定該些觸發訊號的輸出順序,其中該些觸發訊號中的一第一觸發訊號用以觸發該設定電路發出該寫入請求,且該些觸發訊號中的剩餘觸發訊號用以觸發該直接記憶體存取電路發出該至少一讀取請求。 The memory management device of claim 1 further comprises: a control circuit that decodes an instruction from a host processor and determines the dependency of the instruction to set a plurality of trigger signals, and determines the operation status of a plurality of storage spaces in the mapping circuit according to the plurality of external virtual register values corresponding to the trigger signals to set the output sequence of the trigger signals, wherein a first trigger signal among the trigger signals is used to trigger the setting circuit to issue the write request, and the remaining trigger signals among the trigger signals are used to trigger the direct memory access circuit to issue the at least one read request. 一種記憶體管理方法,包含:經由一直接記憶體存取電路取得一原始資料,該原始資料指示一第一虛擬地址與一記憶體的複數個實體地址之間的映射關係;解析該原始資料以將該些實體地址依序映射至包含該第一虛擬地址的複數個第二虛擬地址並發出一寫入請求;根據該寫入請求儲存該些實體地址與該些第二虛擬地址之間的映射關係為一第一映射表,並根據對應於該直接記憶體存取電路的至少一通道之至少一讀取請求利用該第一映射表以存取該記憶體;以及仲裁該寫入請求與該至少一讀取請求,以決定該寫入請求與該至少一讀取請求中每一者存取用來儲存該第一映射表的一儲存空間的順序。 A memory management method comprises: obtaining an original data through a direct memory access circuit, the original data indicating a mapping relationship between a first virtual address and a plurality of physical addresses of a memory; parsing the original data to sequentially map the physical addresses to a plurality of second virtual addresses including the first virtual address and issuing a write request; storing the mapping relationship between the physical addresses and the second virtual addresses according to the write request; The mapping relationship between the second virtual addresses is a first mapping table, and the first mapping table is used to access the memory according to at least one read request corresponding to at least one channel of the direct memory access circuit; and the write request and the at least one read request are arbitrated to determine the order in which each of the write request and the at least one read request accesses a storage space used to store the first mapping table. 一種記憶體管理方法,包含:經由一直接記憶體存取電路取得一原始資料,該原始資料指示一第一虛擬地址與一記憶體的複數個實體地址之間的映射關係;解析該原始資料以將該些實體地址依序映射至包含該第一虛擬地址的複數個第二虛擬地址並發出一寫入請求;以及 根據該寫入請求儲存該些實體地址與該些第二虛擬地址之間的映射關係為一第一映射表,並根據對應於該直接記憶體存取電路的至少一通道之至少一讀取請求利用該第一映射表以存取該記憶體,其中該第一映射表儲存在一映射電路的一第一儲存空間,且當該映射電路欲儲存該第一映射表至該第一儲存空間且該至少一讀取請求欲讀取該映射電路中的一第二儲存空間時,該第一映射表儲存至該第一儲存空間的期間與該第二儲存空間響應於該至少一讀取請求被讀取的期間至少部分重疊。 A memory management method comprises: obtaining an original data through a direct memory access circuit, the original data indicating a mapping relationship between a first virtual address and a plurality of physical addresses of a memory; parsing the original data to sequentially map the physical addresses to a plurality of second virtual addresses including the first virtual address and issuing a write request; and storing the mapping relationship between the physical addresses and the second virtual addresses as a first mapping table according to the write request, and mapping the mapping relationship between the physical addresses and the second virtual addresses according to the first virtual address corresponding to the direct memory access circuit. At least one read request of at least one channel of the memory access circuit uses the first mapping table to access the memory, wherein the first mapping table is stored in a first storage space of a mapping circuit, and when the mapping circuit intends to store the first mapping table in the first storage space and the at least one read request intends to read a second storage space in the mapping circuit, the period during which the first mapping table is stored in the first storage space and the period during which the second storage space is read in response to the at least one read request at least partially overlap.
TW112105770A 2023-02-17 2023-02-17 Memory management device and method applied to intelligence processing unit TWI864600B (en)

Priority Applications (1)

Application Number Priority Date Filing Date Title
TW112105770A TWI864600B (en) 2023-02-17 2023-02-17 Memory management device and method applied to intelligence processing unit

Applications Claiming Priority (1)

Application Number Priority Date Filing Date Title
TW112105770A TWI864600B (en) 2023-02-17 2023-02-17 Memory management device and method applied to intelligence processing unit

Publications (2)

Publication Number Publication Date
TW202435080A TW202435080A (en) 2024-09-01
TWI864600B true TWI864600B (en) 2024-12-01

Family

ID=93609692

Family Applications (1)

Application Number Title Priority Date Filing Date
TW112105770A TWI864600B (en) 2023-02-17 2023-02-17 Memory management device and method applied to intelligence processing unit

Country Status (1)

Country Link
TW (1) TWI864600B (en)

Citations (8)

* Cited by examiner, † Cited by third party
Publication number Priority date Publication date Assignee Title
US20070073996A1 (en) * 2005-04-07 2007-03-29 Ati Technologies Inc. Virtual memory fragment aware cache
US7386697B1 (en) * 2004-01-30 2008-06-10 Nvidia Corporation Memory management for virtual address space with translation units of variable range size
TW201433917A (en) * 2013-01-07 2014-09-01 三星電子股份有限公司 System chip, electronic system and memory address translation method thereof
TWI588654B (en) * 2014-05-09 2017-06-21 美光科技公司 Virtualized physical addresses for reconfigurable memory systems
TW201918882A (en) * 2017-11-13 2019-05-16 韓商愛思開海力士有限公司 Memory system and operating method thereof
CN110659225A (en) * 2018-06-28 2020-01-07 华为技术有限公司 Memory management method and related device
CN111813710A (en) * 2020-09-11 2020-10-23 鹏城实验室 Avoid Linux Kernel Memory Fragmentation Methods, Devices and Computer Storage Media
TW202301127A (en) * 2021-06-17 2023-01-01 日商鎧俠股份有限公司 Memory system and information processing system

Patent Citations (8)

* Cited by examiner, † Cited by third party
Publication number Priority date Publication date Assignee Title
US7386697B1 (en) * 2004-01-30 2008-06-10 Nvidia Corporation Memory management for virtual address space with translation units of variable range size
US20070073996A1 (en) * 2005-04-07 2007-03-29 Ati Technologies Inc. Virtual memory fragment aware cache
TW201433917A (en) * 2013-01-07 2014-09-01 三星電子股份有限公司 System chip, electronic system and memory address translation method thereof
TWI588654B (en) * 2014-05-09 2017-06-21 美光科技公司 Virtualized physical addresses for reconfigurable memory systems
TW201918882A (en) * 2017-11-13 2019-05-16 韓商愛思開海力士有限公司 Memory system and operating method thereof
CN110659225A (en) * 2018-06-28 2020-01-07 华为技术有限公司 Memory management method and related device
CN111813710A (en) * 2020-09-11 2020-10-23 鹏城实验室 Avoid Linux Kernel Memory Fragmentation Methods, Devices and Computer Storage Media
TW202301127A (en) * 2021-06-17 2023-01-01 日商鎧俠股份有限公司 Memory system and information processing system

Also Published As

Publication number Publication date
TW202435080A (en) 2024-09-01

Similar Documents

Publication Publication Date Title
CN107657581B (en) A convolutional neural network CNN hardware accelerator and acceleration method
CN109564545B (en) Method and apparatus for compressing addresses
US20130318285A1 (en) Flash memory controller
JP7630667B2 (en) MEMORY CONTROLLER AND METHODS EMBODIED THEREIN - Patent application
CN108139994B (en) Memory access method and memory controller
CN117312201B (en) A data transmission method, device and accelerator equipment, host and storage medium
KR101789190B1 (en) Cache with scratch pad memory structure and processor including the cache
JP4966404B2 (en) MEMORY CONTROL DEVICE, STORAGE DEVICE, AND MEMORY CONTROL METHOD
JP2013025795A (en) Flash controller hardware architecture for flash devices
CN103077123A (en) Data writing and reading methods and devices
US20050253858A1 (en) Memory control system and method in which prefetch buffers are assigned uniquely to multiple burst streams
CN114168495B (en) Enhanced read-ahead capabilities of storage devices
US7039728B2 (en) Information processing device and method
CN112860596B (en) Data stream cache device of neural network tensor processor
WO2025139618A1 (en) Data reading method for chip, and chip, computer device, storage medium and computer program product
US12436718B2 (en) Memory system
WO2025167292A1 (en) Data read-write method and system, and device and storage medium
US20160110286A1 (en) Data writing method and memory system
CN101261611A (en) Data transmission device and method between peripheral equipment
TWI864600B (en) Memory management device and method applied to intelligence processing unit
US12360929B2 (en) Memory management device and method applied to intelligence processing unit
JP7177948B2 (en) Information processing device and information processing method
CN116107923B (en) BRAM-based many-to-many high-speed memory access architecture and memory access system
CN113157205B (en) Control method of NAND array, controller, electronic device and storage medium
US8484411B1 (en) System and method for improving access efficiency to a dynamic random access memory