US10313108B2 — Energy-efficient bitcoin mining hardware accelerators

Bitcoin Research — Law, Regulation, Markets & Origins (2026)

Patents

2016-06-29

Document text

Research, not advice. Part of the Bitcoin research archive (October 2026). Claims labelled unverified, contested or fringe are reported, not endorsed; statuses of bills and rules are as of the date checked. Government, court and patent records are public domain; the research notes are CC BY 4.0.

US010313108B2

(12) United States Patent                                                                                    ( 10) Patent No.: US 10 ,313 , 108 B2
        Suresh et al.                                                                                        (45 ) Date of Patent:                                        Jun . 4 , 2019
(54 ) ENERGY-EFFICIENT BITCOIN MINING                                                                 (56 )                                       References Cited
        HARDWARE ACCELERATORS
                                                                                                                                         U . S . PATENT DOCUMENTS
(71) Applicant: Intel Corporation, Santa Clara , CA                                                     2004/ 0098630 A15 / 2004 Masleid
                    (US)                                                                                2008/0107259 Al 5 /2008 Satou
                                                                                                        2009/ 0058478 A1 3 / 2009 Lin
                                                                                                        2009 /0119407 Al 5 /2009 Krishnan
(72 ) Inventors: Vikram B . Suresh , Hillsboro , OR                                                     2010 /0086127 A1 4 /2010 Grinchuk et al.
                 (US); Sudhir K . Satpathy, Hillsboro ,                                                 2012 / 0084568 A1                       4 /2012 Sarikaya et al.
                   OR (US ) ; Sanu K . Mathew , Hillsboro ,                                             2013 /0195235 A1                        8/ 2013 Ferris et al.
                   OR (US)                                                                                                                       (Continued )
(73 ) Assignee: Intel Corporation , Santa Clara , CA                                                                             FOREIGN PATENT DOCUMENTS
                    (US )                                                                             WO                          2015077378 A1 5 /2015
                                                                                                      WO                          2016046821 Al 3 /2016
( * ) Notice:      Subject to any disclaimer , the term of this
                   patent is extended or adjusted under 35                                                               OTHER PUBLICATIONS
                    U .S .C . 154 (b ) by 303 days .
                                                                                                      Barkatullah , Goldstrike 1: Coin Terra 's First -Generation Cryptocur
(21) Appl. No.: 15 /196,686                                                                           rency Mining Processor for Bitcoin , 2015 , IEEE ( Year: 2015).*
                                                                                                                            (Continued )
(22) Filed : Jun . 29, 2016                                                                           Primary Examiner — Saleh Najjar
                                                                                                      Assistant Examiner — Louis C Teng
(65 )                  Prior Publication Data                                                         (74) Attorney , Agent, or Firm — Lowenstein Sandler LLP
        US 2018 /0006807 A1          Jan . 4 , 2018                                                   (57 )                 ABSTRACT
                                                                                                       A processing system includes a processor to construct an
(51) Int. Ci.                                                                                         input message comprising a target value and a nonce and a
     H04L 9 /06                  ( 2006 .01)                                                          hardware accelerator, communicatively coupled to the pro
(52 ) U .S . CI.                                                                                      cessor, implementing a plurality of circuits to perform
        CPC ........ H04L 9 / 0643 (2013 .01); H04L 9 / 0675                                          stage - 1 secure hash algorithm (SHA ) hash and stage - 2 SHA
                 (2013.01) ; H04L 2209/ 20 ( 2013 . 01 ); H04L                                        hash , wherein to perform the stage - 2 SHA hash , the hard
               2209/ 30 (2013 .01); H04L 2209/ 38 ( 2013 . 01 ) ;                                     ware accelerator is to perform a plurality of rounds of
                                 H04L 2209 /56 (2013 .01)                                             compression on state data stored in a plurality of registers
(58) Field of Classification Search                                                                   associated with a stage - 2 SHA hash circuit using an input
     CPC . HO4L 9 /0643; H04L 9 /0675 ; H04L 2209 /20 ;                                               value , calculate a plurality of speculative computation bits
                    HO4L 2209 / 30 ; H04L 2209 /38 ; HO4L                                             using a plurality of bits of the state data , and transmit the
                                                                     2209 / 56                        plurality of speculative computation bits to the processor.
        See application file for complete search history.                                                                                20 Claims, 12 Drawing Sheets
                                                                                        400

                                                        Perform stage -O SHA -256 hash and stage - 1 SHA -256 hash on an input
                                                     message comprising a nonce and target value to generate a first hash value
                                                                                      402

                                                         Perform N rounds of stage- 2 ,wherein the N is determined based on a
                                                                      number of speculative computation bits 404

                                                      Store content of state registers associated with the stage - 2 SHA - 256 hash In
                                                                               a second set of registers 406

                                                      Perform 64- N rounds of speculative computation to determine speculative
                                                                                computation bits 408

                                                                                        All SCB = 0
                                                                                           410
                                                Reload the content stored in the                              Update nonce value in
                                                  second set of registers to the                              the input message 412
                                                       state registers 414

                                                   Perform non - speculative
                                                computation in the 64- N rounds
                                                 of stage - 2 SHA - 256 hash 416
                                                          US 10 ,Page
                                                                 313 ,2 108 B2

( 56 )                    References Cited
                 U . S . PATENT DOCUMENTS
 2014 /0093069 A1 *      4 /2014 Wolrich .................. G09C 1/ 00
                                                               380 / 28
 2016 /0109870 A1        4 /2016 Firu et al.
 2016 /0112200 A1        4 / 2016 Kheterpal et al.
 2016 /0125040 A1*       5/ 2016 Kheterpal .......... G06Q 20 / 3678
                                                            707 / 776
 2016 /0164672 Al *      6 / 2016 Karighattam ......... H04L 9/0643
                                                               380 /28
 2018/ 0004242 AL        1/ 2018 Suresh et al.
 2018 / 0006808 A1 *      1/ 2018 Suresh            ... HO4L 9 / 0643
 2018/ 0089642 A1*       3/ 2018 Suresh ............... G06Q 20 /0658
                   OTHER PUBLICATIONS
PCT, International Search Report and Written Opinion of the
International Searching Authority for International Application No .
PCT/US2017 / 035195 dated Sep . 7 , 2017 ( 10 pages).
PCT, International Search Report and Written Opinion of the
International Searching Authority for International Application No .
PCT/US2017 / 035136 , dated Sep . 8 , 2017 , 13 pages.
Dr. Timo Hanke , “ AsicBoost - A Speedup for Bitcoin Mining” ,
Mar. 31 , 2016 (rev .5 ), 10 pages.
 A . Gowthaman et al., “ Performance Study of Enhanced SHA - 256
Algorithm ” , International Journal ofApplied Engineering Research ,
ISSN 0973 - 4562 vol. 10 , No . 4 ( 2015 ), 12 pages.
Robert P . McEvoy et al., “ Optimisation of the SHA- 2 Family of
Hash Functions on FPGAs” , Department of Electrical & Electronic
Engineering, University College Cork , Ireland ,Mar. 2006 , 6 pages .
PCT, International Search Report and Written Opinion of the
International Searching Authority for International Application No.
PCT/US2017/035191 dated Sep . 8, 2017 , 15 pages .
* cited by examiner
U . S . Patent             Jun . 4 , 2019          Sheet 1 of 12                   US 10 ,313, 108 B2

                                      Bus 106
             Bitcoin mining 108                             Stage -O SHA- 256
                                                               engines 110
              Processor 102                                 Stage - 1 SHA - 256
                                                              engines 112
                                                           Speculative stage - ?
                                                           SHA -256 engines 114
                                                                      ASICS 104

                                                  100
                                                Figure 1
U . S . Patent                           Jun . 4 , 2019                     Sheet 2 of 12                            US 10 ,313 ,108 B2

                                                                        200
   218               32                256                       256                 3225632384
                                                                                      32

                 Version 202 Previous Hash 204            Merkel Root 206          Time 208 Target 210 Nonce 212      Bit Padding 214
                                                                                                             41

   ( Constant)                                           512
                       256
                             State                         Message
                                     SHA256 Stage - 0
                                         Output
                                                                                                                   512

                                             First intermediate hash        256
                                                                            .. .
                                                                                                                     Message
                                                                                                    Output
                                                                                                  256      2nd intermediate hash
                                                               (Constant)                                          512             256b Padding
                                                                                   256
                                                                                         State
                                                                                                 SHA256 Stage- 2
                                                                                                    Output

                                 Figure 2                                                        2567
U .S . Patent             Jun . 4 , 2019          Sheet 3 of 12                       US 10 ,313 , 108 B2

           256                             W57            221                                  W57
                               32 0
          State                 Message                   State                      Message
          SHA256 Stage - 2 Round -57                      SHA256 Stage - 2 Round-57
                 Output                                                Output
                  256                      W58                    189                          W58
                                32                                             32
                               Message                    State                Message
          SHA256 Stage- 2 Round -58                        SHA256 Stage - 2 Round -58
                  Output                                               Output
                  256                                             142                          W59
                               32                                               32
          State                Message                     State                Message
          SHA256 Stage - 2 Round - 59                      SHA256 Stage - 2 Round- 59
                    Output                                          Output
                  256                                             16
                                                                                               W50
                                            W60
                               32
          State                Message                                 State Msg
          SHA256 Stage- 2 Round- 60                                    Round- 60
                    Output                                               Output
                   32
                  Figure 3A                                        Figure 3B
U . S . Patent             Jun . 4 , 2019          Sheet 4 of 12              US 10 ,313 ,108 B2

              Perform stage - O SHA - 256 hash and stage - 1 SHA - 256 hash on an input
           message comprising a nonce and target value to generate a first hash value
                                                  402

              Perform N rounds of stage-2 , wherein the N is determined based on a
                          number of speculative computation bits 404

           Store content of state registers associated with the stage-2 SHA -256 hash in
                                   a second set of registers 406

           Perform 64 - N rounds of speculative computation to determine speculative
                                          computation bits 408

                                              All SCB = 0
                                                 410
     Reload the content stored in the                            Update nonce value in
       second set ofregisters to the                             the inputmessage 412
            state registers 414

         Perform non -speculative
      computation in the 64 -N rounds
        of stage - 2 SHA - 256 hash 416                                   Figure 4
U . S . Patent                                                           Jun . 4 , 2019                                                                   Sheet 5 of 12                                   US 10 ,313 ,108 B2

                                  Branch PredictionUnit Instruction cre dunt
                              Branch Predicion Unit                                                                                    Instruction Cache Unit 534
                                                                                                                                        Instruction TLB Unit 536
                                                                                                                                            Instruction Fotch 638
                                         FrontEnd Unit
                                                                                                                                                  Decode Unit 540
                   . ..           .. . . .       .. . . .. .... .. .             .   .. . . . . . .   . .. .. .. . .. .    ...

                              Execution Engine Unit
                                                                       550                                                - ---
                                                                                                                             Rename
                                                                                                                                - - /- -
                                                                                                                                      Allocator
                                                                                                                                           - Unit 552 ]
                                                                       RetirementUnit ; - -- -- -Scheduler Unit( ) 55                                                wwwwwwwwwwwwwwwwww
                                                                              564                 - - . . met
                                                                                                                                 hun wwwwwwwwwwwwwwwwwwwwww

                                                                                       PhysicalRegister Files Unitis )
                                                                             --- - -- - - - - - - - - - - -- - - - - -
                          1-way Hash                                                   Execution Unit(s ) Memory Access
                    Function 590                                                              562                Unit(s ) 564
                                                                                       Execution Cluster(s ) 560
                                                                                                                                                              Data TLB Unit
                                                                                                                          Memory Unit       572                                           L2 Cache Unit
                                                                                                                              570     Data Cache Unit                                         576
                                                                                                                                                                                                          FIG . 5A
                                                                                                             - -                                                                                                       - - - -
                      Length Decode Alloc, Renaming Schedule Register
             Fetch ! Decoding                                         Read Execute Stage Write Back ! Exception
  Pipeline
                                                      512 Memory Read
                                                                                           Memory      Handling Commit
                          -   -
                               506 508 510
                                    -        -
                                                                                                      | 520i 522                              w

                                                                                                                                                                                                            FIG . 5B
U . S . Patent              Jun . 4 , 2019                   Sheet 6 of 12                            US 10 , 313, 108 B2

                                                                        Instruction Prefetcher              Front End
                                                                                  626

                                                                         Instruction Decoder                 Microcode
                                                                                  628                           ROM

                                                                               Trace Cache                   UOP Queue
                                                                                   630
   Processor -
   600                                                     Allocator/Register Renamer
                         Memory UOP                              Integer/Floating Point UOP Queue
   Out Of Order Engine     Queue
   603
                           Memory             Fast Scheduler             Slow /GeneralFP Scheduler            Simple FP
                          Scheduler                 602                                                     Scheduler 606
   Exe Block -
                                      Integer Register File / Bypass Network                     FP Register File / Bypass
                                                                                                       Network 610

                          AGU           AGU          Fast ALU      Fast ALUI Slow ALU                          FP Move
                                                                   1618                            622

                           To Level 1 Cache                                                                 To Level 1 Cache
                                                                 FIG . 6
atent                  Jun . 4 , 2019                    Sheet 7 of 12                    US 10 ,313 ,108 B2

                             Processor 770                         Processor 780

 Memory                                                                             IMC         Memory
    732                    772                                                      782           734

                                                                         786

                                   794
                                                     Chipset 790
 High -Pert Graphics
          738          739
                                 792
                                                                                          716
       BUS Bridge                      1/0 Devices                      Audio i/o
                                                                           724

          Keyboard Mouse                 Comm Devices                                      Data Storage 728
                722                          727
                                                                                            Code And Data
                                                                                                 730
                                                        FIG . 7
us tona mese non consum
U . S . Patent       Jun . 4 , 2019       Sheet 8 of 12         US 10 ,313, 108 B2

                                                   - 815

                    Prd
                                      Processor

          Display                       GMCH                    Memory
            845                          820                     840

                                               - 895

                                         ICH

                     ExternalGraphics              Peripheral
                          Device
                            860

                                        FIG . 8
U . S . Patent   Jun . 4 , 2019                Sheet 9 of 12                     US 10 ,313, 108 B2

                                            10 Devices
                                               914

                      Processor 970                              Processor 980

       Memory                                                                          Memory
        932       %                                                                      934

                                             950
                                                          P -P
                          976         978                              986

                                            Chipset 990

                                                996

                                            Legacy I/O
                                               975

                                            FIG . 9
U . S . Patent                                       Jun . 4 , 2019                                     Sheet 10 of 12                                         US 10 ,313 , 108 B2

          System On A Chip                                                                      Application Processor
              1000                                                                                                   .     .         . ..     . . . ...   W

                                                                   Core 1002A                                            Core 1002N
                                                                                                                                                          WW
                                                                          Cache                                                Cache
                                                                                                                               Unit(s)                    M

                                                                                                                                                                   System Agent Unit
                                                                                                                               1004N
                                                                                                                                                                          1010
                                                                                                                                                          M

                                                                                                                                                          M

                       Media Processors)
                                  1020                                                          Shared Cache Unit(s )
                           - - - - - - - - -
                                                                                                         1006
                       integrated Graphics                  . . . ... .      . . . . . . .. .       .      . .       . . . . . ...      . .. . . .

          -        - - - - - -           - - - - -

                        Image Processor                                                 InterconnectUnit(s) 1002                                                 BUS Controller Unitis )
                              1024                                                                                                                                        1016
      1
      -        - - - - -               - - - - - -

                        Audio Processor
                              1026
      -        -   -          -    -           -

                                                        Integrated Memory                                                                                             Display Unit
                        Video Processor                  Controller Unit(s )                             SRAM Unit                             DMA Unit
                              1028                                                                          1030                                     1032

                                                                                                  FIG . 10
U . S . Patent                         Jun . 4 , 2019                Sheet 11 of 12                             US 10 , 313, 108 B2

                                                                                                     LCD
                                                                                                                       Bluetooth
                                                                                                                          1170
                      Core                  Core
                      1106                   110Z
                                                                    GPU            Video Codec
                         L2 Cache Control 1108                                .

                                                                                          1120   LCD Video /F          3G Modem
                                                                              .

                                                                              .
                                                                                                     1125                1175
              BUS Interface Unit          L2 Cache                            .

                       1109                  1110                             .

                                                                              .

                                                     Interconnect
                                                                                                                         GPS

              SIM             Boot ROM     SORAM Controller               Flash Controller           PC
               1130             1135                1140                          1145               1750
                                                                                                                      802. 11 WiFi
                                                                                                                          1185
                                                   DRAM                           Flash
                                                                                   1165
   Power Control
       1155

                                                            FIG . 11
U .S . Patent             Jun . 4 , 2019          Sheet 12 of 12               US 10 ,313 , 108 B2

                                                                                      1200

       PROCESSOR 1202                                              STATIC MEMORY
                                                                        1206
          PROCESSING
           LOGIC 1226
                                                                   VIDEO DISPLAY
                                                      BUS               1210
      MAIN MEMORY 1204                                1230

          INSTRUCTIONS                                              ALPHA -NUMERIC
                1226                                                 INPUT DEVICE
                                                                         1212

                GRAPHICS
                                           OLUN
            PROCESSING                                                 CURSOR
                  UNIT                                                CONTROL
                   1222                                                DEVICE
                                                                         1214
                 VIDEO
            PROCESSING                                                  SIGNAL
                  UNIT                                               GENERATION
                  1228                                                 DEVICE
                                                                          1216
                 AUDIO
            PROCESSING                                             DATA STORAGE DEVICE
                  UNIT                                                     1218
                   1232
                                                                    MACHINE-READABLE
                                                                       MEDIUM 1224
            NETWORK
            INTERFACE
                DEVICE                                                   SOFTWARE
                  1208                                                         1226

                  NETWORK
                       1220                                               FIG . 12
                                                     US 10 ,313 , 108 B2
       ENERGY-EFFICIENT BITCOIN MINING                                 FIG . 8 is a block diagram of a system in which an
              HARDWARE ACCELERATORS                                  embodiment of the disclosure may operate .
                                                                   FIG . 9 is a block diagram of a system in which an
                     TECHNICAL FIELD                            embodiment of the disclosure may operate .
                                                             5     FIG . 10 is a block diagram of a System -on -a -Chip (SOC )
   The present disclosure relates to hardware accelerators      in accordance with an embodiment of the present disclosure
and , more specifically , to a processing system including a       FIG . 11 is a block diagram of an embodiment of an SoC
processor employing energy - efficient hardware accelerators design in accordance with the present disclosure .
with speculative nonce selection for Bitcoin mining.               FIG . 12 illustrates a block diagram of one embodiment of
                                                             10 a computer system .
                      BACKGROUND                                              DETAILED DESCRIPTION
   Bitcoin is a type of digital currency used in peer-to -peer    The reward for a successful Bitcoin mining is the gen
transactions. The use of Bitcoin in transactions may elimi 15 eration  of a certain number of new Bitcoins ( e .g ., 25
nate the need for intermediate financial institutes because
Bitcoin may enforce authenticity and user anonymity by validated) and
                                                               Bitcoins     the service fee associated with the transactions
                                                                         during the mining process. Each Bitcoin may be
employing digital signatures. Bitcoin resolves the “ double exchanged for currencies in circulation ( e . g ., U . S . dollars ) or
spending” problem (namely , using the same Bitcoin more used in transactions with merchants that accept Bitcoins .
than once by a same entity in different transactions ) using 20 Bitcoin mining may be associated with certain costs such as,
block chaining, whereas a public ledger records all the         for example , the computing resources consumed to perform
transactions that occur within the Bitcoin currency system . Bitcoin mining operations. The most expensive operation in
Every block added to the block chain validates a new set of Bitcoin mining involves the computationally -intensive task
transactions by compressing a 1024 - bit message which               of determining the validity of a 32 -bit nonce . The nonce is
includes a cryptographic root (e. g., the Merkle root) of the 25 a number or a string of bits that is used only once. A 32 -bit
transaction along with bits representing other information nonce is a number (or a string ofbits ) that is represented by
such as , for example , a time stamp associated with the             32 bits. The 32- bit nonce may be part of a 1024 -bit input
transaction, a version number, a target , the hash value of the      message that may also include the Merkle root, the hash of
last block in the block chain and a nonce . The process of the last chain block , and other parameters . The 1024 - bit
validating transactions and generating new blocks of the 30 message may be hashed using three stages of a secure hash
block chain is commonly referred to as Bitcoin mining.      algorithm (e .g., SHA - 256 ) to produce a 256 -bit hash value
                                                            thatmay be compared to a target value also contained in the
      BRIEF DESCRIPTION OF THE DRAWINGS                     input message to determine the validity of the nonce . The
                                                            operations to calculate the hash value are commonly per
   The disclosure will be understood more fully from the 35 formed on hardware accelerators ( e.g., the SHA - 256 hash
detailed description given below and from the accompany .            may be performed on application -specific integrated circuits
ing drawings of various embodiments of the disclosure . The          (ASICs )) and may consume a lot of power . The power
drawings, however , should not be taken to limit the disclo -        consumption by the hardware accelerators is the recurring
sure to the specific embodiments ,but are for explanation and        cost for the Bitcoin mining . Embodiments of the present
understanding only.                                               40 disclosure provide technical solutions including hardware
  FIG . 1 illustrates a processing system to perform Bitcoin         accelerators to perform energy - efficient Bitcoin mining.
mining by employing energy -efficient hardware accelerators             Embodiments of the present disclosure may include
according to an embodiment of the present disclosure .               ASIC - implemented computation of speculative computation
   FIG . 2 illustrates a process to hash a 1024 - bit message        bits that may enable fast identification of an invalid nonce .
into a hash value using three stages of SHA hash in Bitcoin 45 The speculative computation bits are a small amount of
mining                                                               leading bits that can be computed without the determination
   FIG . 3A illustrates rounds 57 -60 of a conventional SHA          of the hash value, thus eliminating the need to compute the
hash .                                                               full 32 bits for potential nonce and reducing the energy
   FIG . 3B illustrates a process to determine the speculative       consumption to perform the full SHA - 256 hash .
computation bits according to an embodiment of the present 50           Bitcoin mining operations include operations to generate
disclosure.                                                          a 256 - bit hash value from a 1024 - bit message . The opera
   FIG . 4 is a block diagram of a method to determine the           tions are part of cryptographic hash that is one -way (very
validity of a nonce using speculative nonce selection in             hard to reverse ) and collision - resistant. The hash operations
stage -2 SHA hash of Bitcoin mining according to an may include two stages ( stage -0 and stage - 1 ) of SHA - 256
embodiment of the present disclosure .                         55 hash to compress a 1024 -bit inputmessage into intermediate
   FIG . 5A is a block diagram illustrating a micro -architec - results , followed by another round (stage - 2 ) of SHA -256
ture for a processor including heterogeneous core in which        hash applied to the intermediate results generated by the first
one embodiment of the disclosure may be used                      two stages of SHA - 256 hash . The 1024 -bit inputmessage to
   FIG . 5B is a block diagram illustrating an in -order pipe- the three stages of SHA - 256 hash contains header informa
line and a register renaming stage , out-of- order issue/ execu - 60 tion , a 32 -bit nonce , and padding bits . The padding bits may
tion pipeline implemented according to at least one embodi-          include 1s and Os that are generated using a padding gen
ment of the disclosure .                                             eration formulae. The 32 -bit nonce is incremented every
   FIG . 6 illustrates a block diagram of the micro -architec        cycle of the Bitcoin mining process to generate an updated
ture for a processor that includes logic in accordance with     input message , where each cycle takes approximate 10
one embodiment of the disclosure.                            65 minutes. A valid nonce is identified if the final hash value
   FIG . 7 is a block diagram illustrating a system in which    contains a certain number of leading zeros. A miner may use
an embodiment of the disclosure may be used .                   the valid nonce as a proof of a successful Bitcoin mining .
                                                  US 10 ,313 , 108 B2
   The software application of Bitcoin miningmay be imple -         hash value may be stored in eight state registers (a , b , c, d ,
mented on a processing system including processors execut-          e, f, g , h ) associated with each SHA - 256 engine, where each
ing Bitcoin mining applications and dedicated hardware
accelerators such as, for examples, ASICs containing clus            word referred to as a state (represented by A , B , C , D , E , F ,
ters of SHA engines that run in parallel to deliver high - 5 G , H ). The initial values of these states can be 32 - bit
performance SHA -256 hash operations . The clusters of SHA   constants . Alternatively , the state registers may initially store
engines may consume a lot of powers (e . g ., at a rate of a hash value calculated from a previous iteration of the
greater than 200 W ) . Embodiments of the present disclosure hashing process . The states ( A , B , C , D , E , F , G , H ) are
include energy -efficient ASIC -based SHA engines that con updated during SHA -256 hash to generate a hash value as
sume less power for Bitcoin mining operations.              10 the
   FIG . 1 illustrates a processing system 100 to perform 10 message
                                                               the output. SHA - 256 hash consumes a block of 512 -bit
Bitcoin mining by employing energy -efficient hardware registers and          compresses it into a 256 -bit hash stored in state
                                                                         (a -h ). The Bitcoin mining process employs three
accelerators including SHA -256 engines according to an
                                                               stages of SHA - 256 hash to convert the 1024 -bit input
embodiment of the present disclosure. As shown in FIG . 1, 15 message
processing system 100 (e.g., a system -on -a-chip (SOC )) 15 me         to a 256 -bit hash value thatmay be compared to a
may include a processor 102 and ASICs 104 communica          target value to determine whether a Bitcoin has been iden
tively coupled to processor 102 via a bus 106 . Processor 102 tified .
may be a hardware processing device such as , for example ,       The SHA -256 hash may include 64 rounds ( identified as
a central processing unit (CPU ) or a graphic processing unit round 0 , 1 , . . . , 63) of applications of compression functions
(GPU ) that includes one or more processing cores (not 20 to the states stored in state registers. The compression
shown ) to execute software applications. Processor 102 may function employs a 512 -bit input value to manipulate the
execute a Bitcoin mining application 108 which may include contents stored in registers (a -h ). Table 1 illustrates the 64
operations to employ multi- stage of SHA - 256 hash to com - rounds of the SHA - 256 operations as applied to the states
press a 1024 -bit inputmessage . For example , Bitcoin mining stored in registers (a -h ) to generate a hash value that can be
application 108 may delegate the calculation of the three 25 used to determine if a valid nonce is found as a proof of the
stages of SHA - 256 hash to hardware accelerators such as,           identification of a Bitcoin .
                                                                                        TABLE 1
                                                        • Apply the SHA -256 compression function to update registers a , b , . . . ,hº
                                                        For j = 0 to 63
                                                          Compute Ch (e, f, g ), Maj(a , b , c), Eo(a ), (e ), and W ; ( see Definitions below )
                                                          T , + h + (e) + Ch(e, f, g) + K ; + W ;
                                                          T2 – Eo (a ) + Maj(a, b , c)
                                                           hag
                                                           gaf
                                                         fce
                                                         ed+ T
                                                          doc
                                                          Cab
                                                          bra
                                                          a = T1 + T2

for example , SHA -256 engines 110 to perform stage -0 hash ,       where logic functions Ch (x , y, z ), Maj(x , y , z ), E , x , & , x are
SHA -256 engines 112 to perform stage - 1 hash , and SHA - 45 compression functions that are defined according the SHA
256 engines 114 to perform stage- 2 hash . These SHA -256           256 specification , and each registers (a -h ) is initiated with a
engines are implemented on one or more ASICS 104 . Each             32 -bit initial values, and W , j = 0 , . . . 63 , are 32 -bit values
one of ASICs 104 may contain multiple SHA -256 engines              derived from a 512 -bit message which can be part of the
( e. g ., > 1000 ) that run in parallel. Embodiments of the    1024 -bit input message of the Bitcoin mining.
present disclosure may take advantage of characteristics of 50 As shown in FIG . 2 , the process of the Bitcoin mining 200
different stages of SHA - 256 hash to implement them in       starts with a 1024 - bit message 218 . The 1024 - bit input
energy efficient manners to save power consumption in               message 218 may be composed of header information , a
Bitcoin mining.                                                     nonce 212 , and padding bits 214 that make the input
   In one embodiment, the stage -2 SHA - 256 hash engine            message 218 to the length of 1024 bits. The header infor
114 (referred to as speculative stage -2 SHA -256 hash 55 mation may include a 32-bit version number 202 , a 256 -bit
engines ) may include speculative nonce selection that may          hash value 204 generated by the immediate preceding block
eliminate a large percentage of invalid nonce based on a             in the block chain of Bitcoin public ledger, a 256 -bit Merkle
small amount of leading bits ( referred to as speculative            root 206 of the transaction , a 32 -bit time stamp 208 , and a
computation bits ) that can be determined quickly without            256 -bit target value 210. Version number 202 is an identifier
incurring the large amount of power consumption to com - 60 associated with the version of the block chain . Hash value
pute the full hash value. Thus, Bitcoin mining application  204 is the hashing result from the immediate preceding
102 may employ speculative SHA- 256 hash engines 114 to block in the block chain recorded in the public ledger.
perform Bitcoin mining at a significantly -reduced rate of Merkle root 206 is the 256 -bit hash based on all of the
power consumption .                                         transactions in the block . Time stamp 208 represents the
  FIG . 2 illustrates a process 200 to hash a 1024 -bit 65 time when the current Bitcoin mining process starts. Target
message into a hash value using three stages of SHA - 256  value 210 represents a threshold value that the resulting hash
hash employed during Bitcoin mining In SHA -256 hash , the           value generated by the Bitcoin mining is compared to . If the
                                                             US 10 ,313 , 108 B2
resulting hash value (“ hash out” ) is smaller than the target                Bitcoin specific optimization. By comparison , both stage- 1
value 210, the nonce 212 in the input message 218 is                          and stage -2 SHA - 256 hash calculations receive inputmes
identified as a valid nonce that can be used as the proof of                  sages relating to the nonce 212 and hence present opportu
the identification of a Bitcoin . If the final result is no less              nities for Bitcoin mining optimizations.
than the target value 210 , the nonce 212 is determined to be 5 In one embodiment of the present disclosure , stage - 2
invalid , or the Bitcoin mining failed to find a Bitcoin . The  SHA - 256 hash may include a speculative nonce pre - selec
value of nonce 212 may be updated ( e . g ., incremented by                   tion that can eliminates a large percentage of invalid nonce
one), and the Bitcoin mining process is repeated to deter      based on the calculation of a few speculative computation
mine the validity of the updated nonce .                       bits . The speculative calculation may calculate the few
   In one embodiment, instead of comparing the final hash - 10 speculative computation bits by performing only part of
ing result with the target value, Bitcoin mining application    stage - 2 SHA - 256 hash . The full stage - 2 SHA - 256 hash is
may determine whether the hash out has a minimum number                       performed only if the speculative calculation cannot deter
of leading zeros. The minimum number of leading zeros                         mine that the nonce is invalid . The quick elimination of a
may ensure that the final hashing value is smaller than the large percentage of nonce can save the power consumption
target value . The target value ( or the number of leading 15 compared to performing the full stage - 2 SHA - 256 hash for
zeros ) may be changed to adjust the complexity of Bitcoin    every nonce .
mining: decreasing the target value decreases the probability   As discussed above , the output of a SHA - 256 hash is
of finding a valid nonce and hence increases the overall commonly a 256 -bit hash value thatmay be stored in eight
search space to generate a new block in the block chain . By
                                                           32 -bit state registers ( a , b , c, d , e , f, g , h ) corresponding to
modifying the target value 210 , the complexity of the 20 eight states ( A , B , C , D , E , F , G , H ). The state registers are
Bitcoin mining is adjusted to ensure that the time used to communicatively accessibly by a processor (e. g., processor
find a valid nonce is relative constant (approximately 10                     102 as shown in FIG . 1 ) executing Bitcoin mining applica
minutes ). For a given header, the Bitcoin mining application                 tion . In one embodiment, the eight states (A , B , C , D , E , F,
may sweep through the search space of 232 possibilities to    G , H ) and their corresponding registers (a , b , c , d , e , f, g , h )
find a valid nonce . The Bitcoin mining process includes a 25 may be arranged in an order from the lowest bit ( e .g ., bits
series of mining iterations to sweeping through these pos -    0 - 31 ) to the highest bits (e . g ., bits 223 - 255 ). Thus , the
sibilities of valid nonce . The header information is kept the                speculative calculation may examine a small number of
same through these mining iterations while the nonce 212 is                   leading bits referred to as the speculative computation bits
incremented by one.                                                           ( e . g ., only the leading two bits ) of state H stored in register
  Each Bitcoin mining calculation to find a valid nonce may 30 h . The number of speculative computation bits is much
include three stages (stage - 0 -stage - 2 ) of SHA - 256 hash smaller than the minimum number of leading zeros required
calculations. Referring to FIG . 2 , at stage -O SHA - 256 hash , to meet the target value 210 . For example, the number of
the state ( A , B , C , D , E , F , G , H ) stored in state registers ( a ,   speculative computation bits may be two bits while the
b , c, d , e , f, g , h ) may be initiated with eight 32 -bit constants .     minimum number of leading zeros required to meet the
Stage- 0 SHA - 256 hash may receive a 512 -bit input message 35 target value 210 is 32 . If any of the speculative computation
including the 32 -bit version number 202 , 256 -bit hash value  bits are non -zero , the nonce can be determined as invalid
204 from the last block in the block chain , and a portion (the without calculating for other bits . Only when all of the
first 224 bits) of Merkle root 206 . Stage -0 SHA - 256 hash                  speculative computation bits are zeros, the nonce can be a
may produce a first 256 -bit intermediate hash value. The first               candidate for a valid nonce and a full stage - 2 SHA - 256 hash
intermediate hash value is then employed to initiate the state 40 is performed to determine if hash outputmeets the minimum
registers A -H of the stage -1 SHA - 256 hash . The 512-bit                   leading zero requirement.
inputmessage to the stage - 1 SHA -256 hash may include the                      The speculative computation bits may help eliminate a
rest portion (32 bits ) of the Merkle root 206 , 32 -bit time                 large portion of invalid nonce without performing the full
stamp 208, 256 -bit target value 210 , 32 -bit nonce 212 , and                stage - 2 SHA -256 hash . For example , when the speculative
128 padding bits 214 . Stage - 1 SHA - 256 hash may produce 45 computation bits are two bits , on average , 75 % of the
a second 256 -bit intermediate hash value .                           candidate nonce can be determined as invalid based on the
    At the stage- 2 SHA - 256 hash , the state registers (a , b , c , two speculative computation bits and be eliminated from
d , e , f, g , h ) of the stage - 2 SHA - 256 hash may be set with    further consideration , thus saving power from performing
the 256 -bit constant same as the constant of stage - 0 SHA -         unnecessary further computation .
256 hash . The 512 -bit inputmessage to the stage - 2 SHA - 256 50 In one embodiment, characteristics of the SHA - 256 hash
hash may include the second 256 -bit intermediate hash                computation may be further explored to reduce the amount
result ( from the stage - 1 SHA - 256 hash output ) combined          of calculation to determine the speculative computation bits .
with 256 padding bits to make a 512 -bit inputmessage to the                  As shown in Table 1, SHA -256 hash includes 64 rounds
stage - 2 SHA - 256 hash . The stage - 2 SHA - 256 hash may                   (rounds 0 -63 ) of applications of compression functions
produce a third 256 -bit hash value as the hash out for the 55 Ch (x , y, z ), Maj (x , y, z ), X , X , Ex to values stored in
three stages of SHA - 256 hash . The Bitcoin mining applica - registers (a , b , c, d , e, f, g , h ) and perform register shift
tion may then determine whether the hash out is smaller than   operations among these registers . Thus , as shown in Table 1 ,
the target value 210 . If the hash out is smaller than the target             the 32 -bit value stored in register h in round 63 is the same
value 210 , the nonce 212 in the input message is identified     as the value stored in register e in round 60 of SHA - 256
as a valid nonce . If the hash out is no less than the target 60 hash . Thus , the leading bits stored for state H can be
value 210 , the nonce 212 is an invalid nonce . After the        determined by the state E in round 60 without performing
determination , nonce 212 is incremented to repeat the pro - the computation of rounds 61 to 63 . Further, the speculative
cess to determine the validity of the updated nonce 212 using                 computation bits may be computed using fewer than the 256
the process as shown in FIG . 2 .                            bits stored in state registers (a , b , c , d , e , f, g , h ).
   Since stage -0 SHA - 256 hash involves only part of the 65 FIG . 3A illustrates rounds 57 -60 of a conventional SHA
header information but not the nonce itself, the calculation                  256 hash . As shown in FIG . 3A , to calculate the 32- bit value
of stage- 0 SHA - 256 does not present an opportunity for                     stored in register e in round 60, each round of the SHA -256
                                                             US 10 ,313, 108 B2
includes the application of compression functions to the                      method 400 may be performed , in part, by processing logics
256 -bit of data stored in registers ( a, b , c, d , e, f, g, h ) using       of processor 102 and ASIC 104 as shown in FIG . 1 .
a 32 -bit word (W ;, j = 57-60 , where j is the round index and                 For simplicity of explanation , the method 400 is depicted
 W , are derived from a 512 -bit input message to stage - 2 and described as a series of acts. However, acts in accor
SHA - 256 hash ) as a key to the compression functions. Thus, 5 dance with this disclosure can occur in various orders and/or
in each round , all 256 bits of data stored in registers (a , b , concurrently and with other acts not presented and described
c , d , e, f, g, h ) are utilized and updated . However, when only            herein . Furthermore, not all illustrated acts may be per
a small number of speculative computation bits (e .g ., two                   formed to implement the method 400 in accordance with the
bits ) are calculated in speculative nonce selection, the num -               disclosed subject matter. In addition , those skilled in the art
ber of bits employed to calculate the speculative computa - 10 will understand and appreciate that the method 400 could
tion bits prior to round 60 can be smaller than 256 bits .     alternatively be represented as a series of interrelated states
   FIG . 3B illustrates a process to determine the speculative via a state diagram or events .
 computation bits according to an embodiment of the present      Referring to FIG . 4 , processor 102 may be communica
disclosure. For the convenience of discussion , it is assumed 15 tively coupled to ASICs 104 which may include clusters of
that two highest bits in the final hash output of stage- 2 SHA -256 engines to perform stage -0 , stage - 1, and stage - 2
SHA -256 hash are employed as the speculative computation SHA - 256 hash for the Bitcoin mining application . Processor
bits . As shown in FIG . 3B , in round 60 , 16 bits ofdata stored 102 may execute a Bitcoin mining application using specu
in registers (a , b , c , d, e, f, g, h ) as the output from round 59, lative computation bits. For the convenience of discussion ,
and two bits of the 32 -bit key W60 are employed to calculate 20 it is assumed that P (e.g., P = 2 ) speculative computation bits
the two speculative computation bits. In round 60 , the 16 bits               are used for the pre -selection , and correspondingly, first N
of data may include W . TO , 1 ), d [0 , 1 ], e [ 0 , 1 , 6 , 7 , 11 , 12 ,   rounds in the stage -2 SHA - 256 hash are performed by using
25 , 26 ], f[0 , 1 ], g [0 , 1 ], h [0 , 1], where X [n , m ] represents X    all of the 256 bits of data stored in stage registers (a, b , c, d ,
register bits at [n , . . . , m ] positions. Similarly, in round 59 ,           e , f, g , h ). The next 64 -N rounds in stage - 2 SHA - 256 may
142 bits of data stored in registers ( a , b , c , d , e , f, g , h ) as the 25 use fewer than 256 bits of data to compute the P speculative
output from round 58 are employed to calculate the 16 bits                      computation bits .
employed in round 60 ; in round 58 , 189 bits of data stored                        At 402 , processor 102 may execute the Bitcoin mining
in registers (a , b , c , d , e, f, g , h ) as the output from round 57       application that employs ASICS 104 to execute stage-0
are employed to calculate the 142 bits employed in round 59;               SHA - 256 hash and stage - 1 SHA - 256 hash on the 1024 - bit
in round 57 , 221 bits of data stored in registers (a , b , c , d , e , 30 input message as discussed in conjunction with FIG . 2 . The
f , g , h ) as the output from round 56 are employed to calculate             input message includes a 32 -bit nonce that needs to be
the 189 bits employed in round 58 . The 221 bits of data                      validated via the three stages of SHA -256 hash . The stage - 1
stored in registers (a , b , c , d , e , f, g , h ) as the output from        SHA -256 hash may generate a first hash value (the second
round 55 may require the full 256 -bit output from round 55 .                 256 -bit intermediate hash value ).
Since the speculative nonce selection is calculated by 35 At 404 , the processor may execute the Bitcoin mining
employing fewer than 256 bits of data stored in (a , b , c , d , application to cause the execution of the first N rounds in
e, f, g , h ), the speculative nonce selection may eliminate a stage -2 SHA -256 hash on ASIC 104 , where the number (N )
large percentage of invalid nonce without incurring the          of rounds is determined based on the number ( P ) of specu
larger power consumption for computing the full stage -2 lative computation bits employed for the speculative nonce
SHA -256 hash .                                               40 selection .
   In one embodiment, responsive to determining that at             At 406 , the processor may store a copy of the contents
least one of the speculative computation bits is non -zero , the              stored in state registers ( a , b , c , d , e , f, g , h ) in a second set
speculative calculation may determine that the nonce 212 in                   of registers (e . g ., a set of temporary registers (a ', b ', c ', d ', e ',
the input message 218 as invalid ; the value of nonce 212 is                  f , g ', h '). The back -up copy of the content stored in state
then updated , and the process to validate nonce is repeated . 45 registers ( a , b , c , d , e , f, g , h , may be used in case that the
Responsive to determining that all of the two speculative                     speculative calculation cannot determine the nonce validity
computation bits are zeros , the nonce 212 in the input                       and the contents in registers ( a , b , c , d , e , f, g , h ) have been
message 218 cannot be determined as invalid based on the                      changed .
speculative computation bits alone . Instead , the rounds ( e .g .,              At 408 , the processor may cause ASICs 104 to perform
rounds 56 -60 when speculative computation bits are two ) of 50 rounds from round 64 - N to round 60 of speculative SHA
SHA - 256 hash are performed to calculate the hash value of 256 hash to determine the P speculative computation bits
the stage - 2 SHA - 256 hash . The hash value generated by that are stored in register e at completion of round 60 . Each
stage -2 SHA -256 hash is then compared to target value 210                   round in round 64 - N to round 60 of the speculative SHA -256
to determine whether nonce 212 contained in inputmessage                      hash employs fewer than 256 bits of content data stored in
218 is valid . If the output hash value is smaller than target 55 registers ( a , b , c , d , e, f, g , h ). In one embodiment, ASIC 104
value 210 , nonce 212 is determined to be valid . If the output               may include circuits that implement these rounds of the
hash value is no less than target value 210 , nonce 212 is                    speculative SHA -256 hash . Thus, processor 102 may
determined to be invalid .                                                    instruct the ASIC 104 to perform these rounds of calculation
   FIG . 4 is a block diagram of a method 400 to determine to generate the P bits of speculative computation bits.
the validity of a nonce using speculative stage - 2 SHA -256 60 At 410 , processor 102 may receive the speculative com
hash of Bitcoin mining according to an embodiment of the                      putation bits and determine whether all of the generated
present disclosure . Method 400 may be performed by pro -                     speculative computation bits are zeros.
cessing logic that may include hardware (e. g ., circuitry,                      If any of the speculative computation bits are non - zero , at
dedicated logic , programmable logic , microcode , etc . ), soft-             412 , processor 102 may determine that the nonce in the input
ware ( such as instructions run on a processing device, a 65 message is invalid , and update the nonce value ( e . g ., incre
general purpose computer system , or a dedicated machine ),                   ment by one) in the 1024 -bit input message to restart the
firmware, or a combination thereof. In one embodiment,                        process to determine the nonce validity .
                                                         US 10 ,313, 108 B2
                                                                                                             10
   If all of the speculative computation bits are zeros,                    illustrate various ways in which register renaming and
processor 102 may determine that the speculative compu                     out -of- order execution may be implemented ( e. g ., using a
tation bits alone cannot determine the validity of the nonce .             reorder buffer (s ) and a retirement register file ( s ), using a
At 414 , processor 102 may reload the contents stored in the                future file (s ), a history buffer( s), and a retirement register
second set of registers ( a ', b ', c ', d ', e', f', g ', h ') back into the 5 file (s ); using a register maps and a pool of registers ; etc .).
state registers (a , b , c, d , e , f, g , h ).                                     In one implementation , processor 500 may be the sameas
   At 416 , processor 102 may cause ASICS 104 to perform                        processor 102 described with respect to FIG . 1 .
rounds 64 - N to 60 in a non - speculative manner to determine                Generally , the architectural registers are visible from the
the leading bits of the final hash out for the stage - 2 SHA -256 outside of the processor or from a programmer 's perspec
hash . After round 60, the leading bits of the hash value 10 tive . The registers are not limited to any known particular
 generated by the stage - 2 SHA - 256 may be stored in stage      type of circuit . Various different types of registers are
register 3 and may be compared to the target value (e . g ., by   suitable as long as they are capable of storing and providing
counting leading zeros ) to determine whether the nonce in                  data as described herein . Examples of suitable registers
the input message is valid . A valid nonce can be used as                  include, but are not limited to , dedicated physical registers ,
proof of a successful Bitcoin mining . The nonce in the input 15 dynamically allocated physical registers using register
message may be updated to search for a next valid nonce .                   renaming, combinations of dedicated and dynamically allo
   FIG . 5A is a block diagram illustrating a micro -architec -             cated physical registers, etc . The retirement unit 554 and the
ture for a processor 500 that implements the processing                    physical register file ( s ) unit( s ) 558 are coupled to the execu
device including heterogeneous cores in accordance with                    tion cluster (s ) 560. The execution cluster ( s ) 560 includes a
one embodiment of the disclosure . Specifically , processor 20 set of one or more execution units 562 and a set of one or
500 depicts an in - order architecture core and a register                 more memory access units 564. The execution units 562
renaming logic , out-of- order issue / execution logic to be               may perform various operations ( e . g ., shifts , addition , sub
included in a processor according to at least one embodi-                  traction , multiplication ) and operate on various types of data
ment of the disclosure .                                       (e.g ., scalar floating point, packed integer, packed floating
  Processor 500 includes a front end unit 530 coupled to an 25 point, vector integer, vector floating point).
execution engine unit 550 , and both are coupled to a                         While some embodiments may include a number of
memory unit 570 . The processor 500 may include a reduced       execution units dedicated to specific functions or sets of
instruction set computing (RISC ) core , a complex instruc -    functions, other embodiments may include only one execu
tion set computing (CISC ) core , a very long instruction word tion unit or multiple execution units that all perform all
(VLIW ) core , or a hybrid or alternative core type . As yet 30 functions. The scheduler unit(s ) 556 , physical register file (s )
another option , processor 500 may include a special-purpose               unit(s ) 558 , and execution cluster(s ) 560 are shown as being
core , such as, for example , a network or communication                   possibly plural because certain embodiments create separate
core, compression engine , graphics core , or the like. In one             pipelines for certain types of data /operations (e.g ., a scalar
embodiment, processor 500 may be a multi-core processor                    integer pipeline, a scalar floating point/ packed integer /
or may part of a multi-processor system .                              35 packed floating point/vector integer /vector floating point
   The front end unit 530 includes a branch prediction unit                pipeline , and/or a memory access pipeline that each have
532 coupled to an instruction cache unit 534 , which is                     their own scheduler unit, physical register file( s) unit , and /or
coupled to an instruction translation lookaside buffer ( TLB )              execution cluster — and in the case of a separate memory
536 , which is coupled to an instruction fetch unit 538 , which             access pipeline, certain embodiments are implemented in
is coupled to a decode unit 540 . The decode unit 540 ( also 40 which only the execution cluster of this pipeline has the
known as a decoder ) may decode instructions , and generate                memory access unit( s ) 564) . It should also be understood
as an output one or more micro -operations,micro - code entry              that where separate pipelines are used , one or more of these
points ,microinstructions, other instructions , or other control pipelines may be out -of -order issue / execution and the rest
signals, which are decoded from , or which otherwise reflect, in -order.
or are derived from , the original instructions . The decoder 45 The set of memory access units 564 is coupled to the
540 may be implemented using various different mecha - memory unit 570 , which may include a data prefetcher 580 ,
nisms. Examples of suitable mechanismsinclude , but are not a data TLB unit 572 , a data cache unit (DCU ) 574 , and a
limited to , look -up tables, hardware implementations , pro     level 2 (L2) cache unit 576 , to name a few examples. In
grammable logic arrays (PLAS ), microcode read only              some embodiments DCU 574 is also known as a first level
memories (ROMs), etc . The instruction cache unit 534 is 50 data cache (L1 cache ). The DCU 574 may handle multiple
further coupled to thememory unit 570. The decode unit 540       outstanding cache misses and continue to service incoming
is coupled to a rename/allocator unit 552 in the execution       stores and loads. It also supports maintaining cache coher
 engine unit 550.                                                ency . The data TLB unit 572 is a cache used to improve
   The execution engine unit 550 includes the renamel                       virtual address translation speed by mapping virtual and
allocator unit 552 coupled to a retirement unit 554 and a set 55 physical address spaces . In one exemplary embodiment, the
of one or more scheduler unit ( s ) 556 . The scheduler unit (s ) memory access units 564 may include a load unit , a store
556 represents any number of different schedulers , including               address unit , and a store data unit, each of which is coupled
reservations stations (RS ), central instruction window , etc .            to the data TLB unit 572 in the memory unit 570 . The L2
The scheduler unit( s ) 556 is coupled to the physical register             cache unit 576 may be coupled to one or more other levels
file ( s ) unit ( s ) 558 . Each of the physical register file (s ) units 60 of cache and eventually to a main memory .
558 represents one or more physical register files, different                  In one embodiment, the data prefetcher 580 speculatively
ones of which store one or more different data types , such as         loads /prefetches data to the DCU 574 by automatically
scalar integer , scalar floating point, packed integer, packed         predicting which data a program is about to consume.
floating point, vector integer, vector floating point, etc .,          Prefeteching may refer to transferring data stored in one
status ( e . g ., an instruction pointer that is the address of the 65 memory location of a memory hierarchy ( e. g ., lower level
next instruction to be executed ), etc. The physical register caches or memory ) to a higher-levelmemory location that is
file (s) unit(s) 558 is overlapped by the retirement unit 554 to            closer (e.g ., yields lower access latency ) to the processor
                                                        US 10 ,313 , 108 B2
                                                                                                       12
before the data is actually demanded by the processor.More               machine can execute . In other embodiments, the decoder
specifically, prefetching may refer to the early retrieval of            parses the instruction into an opcode and corresponding data
data from one of the lower level caches/memory to a data                 and control fields that are used by the micro - architecture to
cache and /or prefetch buffer before the processor issues a              perform operations in accordance with one embodiment. In
demand for the specific data being returned .                          5 one embodiment, the trace cache 630 takes decoded uops
   The processor 500 may support one or more instructions                and assembles them into program ordered sequences or
sets ( e . g., the x86 instruction set (with some extensions that        traces in the uop queue 634 for execution . When the trace
have been added with newer versions ); the MIPS instruction              cache 630 encounters a complex instruction , the microcode
set of MIPS Technologies of Sunnyvale , Calif.; the ARM                  ROM 632 provides the uops needed to complete the opera
instruction set (with optional additional extensions such as 10 tion .
NEON ) of ARM Holdings of Sunnyvale, Calif .).                     Some instructions are converted into a single micro -op ,
   It should be understood that the core may support multi-              whereas others need several micro -ops to complete the full
threading (executing two or more parallel sets of operations             operation . In one embodiment, if more than four micro -ops
or threads ), and may do so in a variety of ways including are needed to complete an instruction , the decoder 628
time sliced multithreading, simultaneous multithreading 15 accesses the microcode ROM 632 to do the instruction . For
(where a single physical core provides a logical core for each           one embodiment, an instruction can be decoded into a small
of the threads that physical core is simultaneously multi-               number of micro ops for processing at the instruction
threading ), or a combination thereof ( e . g ., time sliced fetch -     decoder 628 . In another embodiment, an instruction can be
ing and decoding and simultaneous multithreading thereaf-     stored within the microcode ROM 632 should a number of
ter such as in the Intel® Hyperthreading technology ) .   20 micro - ops be needed to accomplish the operation . The trace
   While register renaming is described in the context of                cache 630 refers to an entry point programmable logic array
out-of-order execution , it should be understood that register           (PLA ) to determine a correct micro - instruction pointer for
renaming may be used in an in -order architecture . While the            reading the micro - code sequences to complete one or more
illustrated embodiment of the processor also includes a      instructions in accordance with one embodiment from the
separate instruction and data cache units and a shared L2 25 micro -code ROM 632 . After the microcode ROM 632
cache unit, alternative embodiments may have a single         finishes sequencing micro -ops for an instruction , the front
internal cache for both instructions and data, such as, for              end 601 of the machine resumes fetching micro -ops from the
example , a Level 1 (L1) internal cache , or multiple levels of          trace cache 630 .
internal cache. In some embodiments, the system may                         The out- of -order execution engine 603 is where the
include a combination of an internal cache and an external 30 instructions are prepared for execution . The out-of- order
cache that is external to the core and /or the processor.                execution logic has a number of buffers to smooth out and
Alternatively, all of the cache may be external to the core              re -order the flow of instructions to optimize performance as
and / or the processor.                                                  they go down the pipeline and get scheduled for execution .
   FIG . 5B is a block diagram illustrating an in -order pipe        The allocator logic allocates the machine buffers and
line and a register renaming stage , out-of- order issue/ execu - 35 resources that each uop needs in order to execute. The
tion pipeline implemented by processing device 500 of FIG . register renaming logic renames logic registers onto entries
5A according to some embodiments of the disclosure . The                 in a register file . The allocator also allocates an entry for
solid lined boxes in FIG . 5B illustrate an in -order pipeline ,         each uop in one of the two uop queues, one for memory
while the dashed lined boxes illustrates a register renaming,            operations and one for non -memory operations, in front of
out-of-order issuelexecution pipeline. In FIG . 5B , a proces- 40 the instruction schedulers : memory scheduler, fast scheduler
sor pipeline 500 includes a fetch stage 502 , a length decode     602 , slow /general floating point scheduler 604 , and simple
stage 504 , a decode stage 506 , an allocation stage 508 , a              floating point scheduler 606 . The uop schedulers 602, 604 ,
renaming stage 510 , a scheduling ( also known as a dispatch             606 , determine when a uop is ready to execute based on the
or issue) stage 512 , a register read /memory read stage 514 ,           readiness of their dependent input register operand sources
an execute stage 516 , a write back /memory write stage 518 , 45 and the availability of the execution resources the uops need
an exception handling stage 522 , and a commit stage 524 . In    to complete their operation . The fast scheduler 602 of one
some embodiments , the ordering of stages 502 -524 may be                embodiment can schedule on each half of the main clock
different than illustrated and are not limited to the specific           cycle while the other schedulers can only schedule once per
ordering shown in FIG . 5B .                                             main processor clock cycle . The schedulers arbitrate for the
   FIG . 6 illustrates a block diagram of the micro - architec - 50 dispatch ports to schedule uops for execution .
ture for a processor 600 that includes hybrid cores in                     Register files 608 , 610 , sit between the schedulers 602 ,
accordance with one embodiment of the disclosure. In some                604, 606, and the execution units 612 , 614 , 616 , 618 , 620 ,
embodiments, an instruction in accordance with one                       622, 624 in the execution block 611 . There is a separate
embodiment can be implemented to operate on data ele                     register file 608 , 610 , for integer and floating point opera
ments having sizes of byte , word , doubleword , quadword , 55 tions, respectively . Each register file 608 , 610, of one
etc ., as well as datatypes, such as single and double precision         embodiment also includes a bypass network that can bypass
integer and floating point datatypes . In one embodiment the or forward just completed results that have not yet been
in -order front end 601 is the part of the processor 600 that    written into the register file to new dependent uops. The
fetches instructions to be executed and prepares them to be integer register file 608 and the floating point register file
used later in the processor pipeline.                         60 610 are also capable of communicating data with the other.
    The front end 601 may include several units . In one For one embodiment, the integer register file 608 is split into
embodiment, the instruction prefetcher 626 fetches instruc -             two separate register files, one register file for the low order
tions from memory and feeds them to an instruction decoder               32 bits of data and a second register file for the high order
628 which in turn decodes or interprets them . For example ,             32 bits of data . The floating point register file 610 of one
in one embodiment, the decoder decodes a received instruc- 65 embodimenthas 128 bit wide entries because floating point
tion into one or more operations called “ micro -instructions” instructions typically have operands from 64 to 128 bits in
or “ micro -operations ” (also called micro op or uops ) that the        width .
                                                   US 10 ,313, 108 B2
                              13                                                                    14
   The execution block 611 contains the execution units 612,         register renaming, combinations of dedicated and dynami
614 , 616 , 618 , 620 , 622 , 624 , where the instructions are       Call
                                                                     cally allocated physical registers , etc . In one embodiment,
actually executed . This section includes the register files         integer registers store thirty -two bit integer data . A register
608 , 610 , that store the integer and floating point data           file of one embodiment also contains eight multimedia
operand values that the micro -instructions need to execute . 5 SIMD registers for packed data .
The processor 600 of one embodiment is comprised of a             For the discussions below , the registers are understood to
number of execution units : address generation unit ( AGU )          be data registers designed to hold packed data , such as 64
612, AGU 614 , fast ALU 616 , fast ALU 618 , slow ALU 620,           bits wide MMXTM registers (also referred to as ‘mm ’ reg
floating point ALU 622 , floating point move unit 624 . For          isters in some instances ) in microprocessors enabled with
one embodiment , the floating point execution blocks 622 , 10 MMX technology from Intel Corporation of Santa Clara ,
624 , execute floating point, MMX , SIMD, and SSE , or other         Calif. These MMX registers, available in both integer and
operations . The floating point ALU 622 of one embodiment            floating point forms, can operate with packed data elements
includes a 64 bit by 64 bit floating point divider to execute        that accompany SIMD and SSE instructions. Similarly , 128
divide , square root, and remainder micro - ops . For embodi-        bits wide XMM registers relating to SSE2 , SSE3 , SSE4 , or
ments of the present disclosure, instructions involving a 15 beyond ( referred to generically as “ SSEX ” ) technology can
floating point value may be handled with the floating point          also be used to hold such packed data operands. In one
hardware .                                                           embodiment, in storing packed data and integer data, the
   In one embodiment, the ALU operations go to the high -            registers do not need to differentiate between the two data
speed ALU execution units 616 , 618 . The fast ALUS 616 , types . In one embodiment, integer and floating point are
618 , of one embodiment can execute fast operations with an 20 either contained in the same register file or different register
effective latency of half a clock cycle . For one embodiment,  files. Furthermore , in one embodiment, floating point and
most complex integer operations go to the slow ALU 620 as integer data may be stored in different registers or the same
the slow ALU 620 includes integer execution hardware for registers.
long latency type of operations, such as a multiplier, shifts,          Referring now to FIG . 7 , shown is a block diagram
flag logic , and branch processing . Memory load / store opera - 25 illustrating a system 700 in which an embodiment of the
tions are executed by the AGUS 612 , 614 . For one embodi-           disclosure may be used . As shown in FIG . 7 , multiprocessor
ment, the integer ALUS 616 , 618 , 620 , are described in the        system 700 is a point-to -point interconnect system , and
context of performing integer operations on 64 bit data              includes a first processor 770 and a second processor 780
operands. In alternative embodiments , the ALUS 616 , 618 ,          coupled via a point -to -point interconnect 750 . While shown
620 , can be implemented to support a variety of data bits 30 with only two processors 770 , 780 , it is to be understood that
including 16 , 32 , 128 , 256 , etc . Similarly , the floating point the scope of embodiments of the disclosure is not so limited .
units 622 , 624 , can be implemented to support a range of           In other embodiments , one or more additional processors
operands having bits of various widths . For one embodi-             may be present in a given processor. In one embodiment, the
ment, the floating point units 622 , 624 , can operate on 128        multiprocessor system 700 may implement hybrid cores as
bits wide packed data operands in conjunction with SIMD 35 described herein .
and multimedia instructions.                                     Processors 770 and 780 are shown including integrated
   In one embodiment, the uops schedulers 602 , 604 , 606 , memory controller units 772 and 782, respectively . Proces
dispatch dependent operations before the parent load has sor 770 also includes as part of its bus controller units
finished executing. As uops are speculatively scheduled and point-to -point (P -P ) interfaces 776 and 778 ; similarly , sec
executed in processor 600 , the processor 600 also includes 40 ond processor 780 includes P -P interfaces 786 and 788 .
logic to handle memory misses. If a data load misses in the          Processors 770 , 780 may exchange information via a point
data cache , there can be dependent operations in flight in the      to - point (PPP ) interface 750 using P -Pinterface circuits 778 ,
pipeline that have left the scheduler with temporarily incor         788 . As shown in FIG . 7, IMCs 772 and 782 couple the
rect data . A replay mechanism tracks and re - executes              processors to respective memories , namely a memory 732
instructions that use incorrect data . Only the dependent 45 and a memory 734 , which may be portions ofmain memory
operations need to be replayed and the independent ones are          locally attached to the respective processors .
allowed to complete . The schedulers and replay mechanism              Processors 770 , 780 may each exchange information with
of one embodiment of a processor are also designed to catch          a chipset 790 via individual P - P interfaces 752, 754 using
instruction sequences for text string comparison operations.         point to point interface circuits 776 , 794 , 786 , 798 . Chipset
   The processor 600 also includes logic to implement store 50 790 may also exchange information with a high -perfor
address prediction formemory disambiguation according to mance graphics circuit 738 via a high - performance graphics
embodiments of the disclosure . In one embodiment, the interface 739 .
execution block 611 of processor 600 may include a store                   shared cache (not shown ) may be included in either
address predictor (not shown ) for implementing store                processor or outside of both processors, yet connected with
address prediction for memory disambiguation .                    55 the processors via P -P interconnect, such that either or both
   The term “ registers ” may refer to the on -board processor       processors ' local cache information may be stored in the
storage locations that are used as part of instructions to           shared cache if a processor is placed into a low power mode.
identify operands. In other words, registers may be those               Chipset 790 may be coupled to a first bus 716 via an
that are usable from the outside of the processor (from a            interface 796 . In one embodiment, first bus 716 may be a
programmer ' s perspective ). However, the registers of an 60 Peripheral Component Interconnect (PCI) bus , or a bus such
embodiment should not be limited in meaning to a particular          as a PCI Express bus or another third generation I/ O
type of circuit. Rather, a register of an embodiment is        interconnect bus, although the scope of the present disclo
capable of storing and providing data , and performing the     sure is not so limited .
functions described herein . The registers described herein       As shown in FIG . 7 , various I/ O devices 714 may be
can be implemented by circuitry within a processor using 65 coupled to first bus 716 , along with a bus bridge 718 which
any number of different techniques , such as dedicated physi-  couples first bus 716 to a second bus 720 . In one embodi
cal registers , dynamically allocated physical registers using ment, second bus 720 may be a low pin count (LPC ) bus.
                                                   US 10 ,313,108 B2
                             15                                                                      16
Various devices may be coupled to second bus 720 includ -            point-to - point interconnect 950 between point-to -point ( P - P )
ing, for example, a keyboard and / or mouse 722, communi             interfaces 978 and 988 respectively . Processors 970, 980
cation devices 727 and a storage unit 728 such as a disk drive       each communicate with chipset 990 via point- to - point inter
or other mass storage device which may include instruc               connects 952 and 954 through the respective P -P interfaces
tions/code and data 730 , in one embodiment . Further, an 5 976 to 994 and 986 to 998 as shown. For at least one
audio I/ O 724 may be coupled to second bus 720 . Note that          embodiment, the CL 972, 982 may include integrated
other architectures are possible . For example , instead of the      memory controller units . CLs 972 , 982 may include I/ O
point-to -point architecture of FIG . 7 , a system may imple -       control logic . As depicted , memories 932, 934 coupled to
ment a multi-drop bus or other such architecture .                   CLs 972 , 982 and I/ O devices 914 are also coupled to the
  Referring now to FIG . 8 , shown is a block diagram of a 10 control logic 972 , 982 . Legacy I/ O devices 915 are coupled
system 800 in which one embodiment of the disclosure may             to the chipset 990 via interface 996 .
operate . The system 800 may include one or more proces -            Embodiments may be implemented in many different
sors 810 , 815 , which are coupled to graphics memory             system types . FIG . 10 is a block diagram of a SoC 1000 in
controller hub (GMCH ) 820 . The optional nature of addi - accordance with an embodiment of the present disclosure .
tional processors 815 is denoted in FIG . 8 with broken lines. 15 Dashed lined boxes are optional features on more advanced
In one embodiment, processors 810 , 815 implementhybrid               SoCs. In FIG . 10 , an interconnect unit(s ) 1012 is coupled to :
cores according to embodiments of the disclosure .                   an application processor 1020 which includes a set of one or
   Each processor 810 , 815 may be some version of the               more cores 1002A -N and shared cache unit(s ) 1006 ; a
circuit , integrated circuit , processor, and / or silicon inte      system agent unit 1010 ; a bus controller unit(s ) 1016 ; an
grated circuit as described above . However , it should be 20 integrated memory controller unit( s ) 1014 ; a set or one or
noted that it is unlikely that integrated graphics logic and more media processors 1018 which may include integrated
integrated memory control units would exist in the proces -   graphics logic 1008 , an image processor 1024 for providing
sors 810 , 815 . FIG . 8 illustrates that the GMCH 820 may be         still and /or video camera functionality , an audio processor
coupled to a memory 840 that may be, for example , a    1026 for providing hardware audio acceleration , and a video
dynamic random access memory (DRAM ). The DRAM 25 processor 1028 for providing video encode /decode accel
may, for at least one embodiment, be associated with a eration ; an static random access memory (SRAM ) unit 1030 ;
non - volatile cache.                                  a directmemory access (DMA ) unit 1032 ; and a display unit
   The GMCH 820 may be a chipset, or a portion of a                  1040 for coupling to one or more external displays. In one
chipset. The GMCH 820 may communicate with the pro                   embodiment, a memory module may be included in the
cessor (s ) 810 , 815 and control interaction between the 30 integrated memory controller unit( s ) 1014 . In another
processor( s ) 810 , 815 and memory 840 . The GMCH 820               embodiment, the memory module may be included in one or
may also act as an accelerated bus interface between the             more other components of the SoC 1000 that may be used
processor( s ) 810 , 815 and other elements of the system 800 .      to access and / or control a memory. The application proces
For at least one embodiment, theGMCH 820 communicates                sor 1020 may include a store address predictor for imple
with the processor ( s ) 810 , 815 via a multi -drop bus, such as 35 menting hybrid cores as described in embodiments herein .
a frontside bus ( FSB ) 895 .                                           The memory hierarchy includes one or more levels of
  Furthermore , GMCH 820 is coupled to a display 845                 cache within the cores, a set or one or more shared cache
(such as a flat panel or touchscreen display ) . GMCH 820            units 1006 , and externalmemory (not shown ) coupled to the
may include an integrated graphics accelerator. GMCH 820             set of integrated memory controller units 1014 . The set of
is further coupled to an input/ output (1 / 0 ) controller hub 40 shared cache units 1006 may include one or more mid -level
( ICH ) 850 , which may be used to couple various peripheral         caches , such as level 2 (L2 ), level 3 (L3 ), level 4 (L4 ), or
devices to system 800 . Shown for example in the embodi-             other levels of cache, a last level cache (LLC ), and/ or
ment of FIG . 8 is an external graphics device 860, which            combinations thereof.
may be a discrete graphics device , coupled to ICH 850 ,            In some embodiments , one or more of the cores 1002A - N
along with another peripheral device 870 .                    45 are capable of multi-threading. The system agent 1010
  Alternatively , additional or different processors may also    includes those components coordinating and operating cores
be present in the system 800. For example, additional                1002A -N . The system agent unit 1010 may include for
processor (s ) 815 may include additional processors ( s ) that   example a power control unit (PCU ) and a display unit . The
are the same as processor 810 , additional processor( s ) that PCU may be or include logic and components needed for
are heterogeneous or asymmetric to processor 810 , accel- 50 regulating the power state of the cores 1002A - N and the
erators (such as, e .g ., graphics accelerators or digital signal integrated graphics logic 1008 . The display unit is for
processing (DSP ) units ), field programmable gate arrays , or       driving one or more externally connected displays .
any other processor. There can be a variety of differences          The cores 1002A - N may be homogenous or heteroge
between the processor ( s ) 810 , 815 in terms of a spectrum of neous in terms of architecture and / or instruction set. For
metrics of merit including architectural,micro -architectural, 55 example , some of the cores 1002A - N may be in order while
thermal, power consumption characteristics, and the like .           others are out -of-order. As another example , two or more of
These differences may effectively manifest themselves as             the cores 1002A -N may be capable of execution the same
asymmetry and heterogeneity amongst the processors 810 ,             instruction set , while others may be capable of executing
815 . For at least one embodiment, the various processors            only a subset of that instruction set or a different instruction
810 , 815 may reside in the same die package .                    60 set.
  Referring now to FIG . 9 , shown is a block diagram of a              The application processor 1020 may be a general- purpose
system 900 in which an embodiment of the disclosure may processor, such as a CoreTM i3 , i5 , i7, 2 Duo and Quad ,
operate . FIG . 9 illustrates processors 970 , 980 . In one XeonTM , ItaniumTM , AtomTM or QuarkTM processor, which
embodiment, processors 970 , 980 may implement hybrid          are available from IntelTM Corporation , ofSanta Clara , Calif.
cores as described above . Processors 970 , 980 may include 65 Alternatively, the application processor 1020 may be from
integrated memory and I/ O control logic (“ CL ” ) 972 and     another company, such as ARM HoldingsTM , Ltd , MIPSTM ,
982, respectively and intercommunicate with each other via etc . The application processor 1020 may be a special
                                                    US 10 ,313 , 108 B2
                          17                                                                       18
purpose processor, such as, for example , a network or              multiple sets) of instructions to perform any one or more of
communication processor, compression engine , graphics              the methodologies discussed herein .
processor, co -processor, embedded processor, or the like.             The computer system 1200 includes a processing device
The application processor 1020 may be implemented on one            1202 , a main memory 1204 ( e. g., read -only memory
or more chips . The application processor 1020 may be a part 5 (ROM ), flash memory , dynamic random access memory
of and /or may be implemented on one or more substrates (DRAM ) (such as synchronous DRAM (SDRAM ) or
using any of a number of process technologies , such as , for DRAM (RDRAM ), etc .), a static memory 1206 ( e .g ., flash
example, BiCMOS, CMOS , or NMOS .
   FIG . 11 is a block diagram of an embodimentof a system memory         , static random access memory (SRAM ), etc.), and
                                                                  a data storage device 1218 , which communicate with each
on -chip (SoC ) design in accordance with the present disclo - 10 other via a bus 1230.
sure. As a specific illustrative example , SoC 1100 is included      Processing device 1202 represents one or more general
in user equipment (UE). In one embodiment, UE refers to purpose            processing devices such as a microprocessor, cen
any device to be used by an end -user to communicate , such
as a hand -held phone , smartphone, tablet, ultra -thin note      tral processing   unit, or the like. More particularly, the
book , notebook with broadband adapter, or any other similar 15 processing device may be complex instruction set comput
communication device . Often a UE connects to a base                ing (CISC ) microprocessor, reduced instruction set com
station or node, which potentially corresponds in nature to a       puter (RISC ) microprocessor, very long instruction word
mobile station (MS) in a GSM network .                              (VLIW ) microprocessor, or processor implementing other
  Here , SOC 1100 includes 2 cores — 1106 and 1107. Cores           instruction sets, or processors implementing a combination
1106 and 1107 may conform to an Instruction Set Architec - 20 of instruction sets . Processing device 1202 may also be one
ture , such as an Intel® Architecture CoreTM -based processor,      or more special -purpose processing devices such as an
an Advanced Micro Devices, Inc . (AMD ) processor, a                application specific integrated circuit (ASIC ) , a field pro
MIPS -based processor, an ARM -based processor design , or          grammable gate array (FPGA ), a digital signal processor
a customer thereof, as well as their licensees or adopters .        (DSP ), network processor, or the like . In one embodiment,
Cores 1106 and 1107 are coupled to cache control 1108 that 25 processing device 1202 may include one or processing
is associated with bus interface unit 1109 and L2 cache 1110        cores. The processing device 1202 is configured to execute
to communicate with other parts of system 1100 . Intercon -         the processing logic 1226 for performing the operations and
nect 1110 includes an on - chip interconnect, such as an IOSF ,     steps discussed herein . In one embodiment, processing
AMBA , or other interconnect discussed above, which poten -         device 1202 is the same as processor architecture 100
tially implements one or more aspects of the described 30 described with respect to FIG . 1 as described herein with
disclosure . In one embodiment, cores 1106 , 1107 may embodiments of the disclosure .
implement hybrid cores as described in embodiments herein .  The computer system 1200 may further include a network
   Interconnect 1110 provides communication channels to             interface device 1208 communicably coupled to a network
the other components , such as a Subscriber Identity Module         1220 . The computer system 1200 also may include a video
( SIM ) 1130 to interface with a SIM card , a boot ROM 1135 35 display unit 1210 (e.g ., a liquid crystal display (LCD ) or a
to hold boot code for execution by cores 1106 and 1107 to            cathode ray tube (CRT )), an alphanumeric input device 1212
initialize and boot SoC 1100 , a SDRAM controller 1140 to (e.g., a keyboard ), a cursor control device 1214 (e.g., a
interface with externalmemory (e.g . DRAM 1160 ), a flash mouse ), and a signal generation device 1216 (e. g., a
controller 1145 to interface with non - volatile memory ( e.g . speaker ). Furthermore , computer system 1200 may include
Flash 1165 ), a peripheral control 1150 (e .g . Serial Peripheral 40 a graphics processing unit 1222 , a video processing unit
Interface ) to interface with peripherals, video codecs 1120         1228 , and an audio processing unit 1232 .
and Video interface 1125 to display and receive input (e .g .           The data storage device 1218 may include a machine
touch enabled input), GPU 1115 to perform graphics related accessible storage medium 1224 on which is stored software
computations, etc . Any of these interfaces may incorporate 1226 implementing any one ormore of the methodologies of
aspects of the disclosure described herein . In addition , the 45 functions described herein , such as implementing store
system 1100 illustrates peripherals for communication , such        address prediction for memory disambiguation as described
as a Bluetooth module 1170 , 3G modem 1175 , GPS 1180,              above. The software 1226 may also reside , completely or at
and Wi-Fi 1185 .                                                    least partially , within the main memory 1204 as instructions
   FIG . 12 illustrates a diagrammatic representation of a          1226 and /or within the processing device 1202 as processing
machine in the example form of a computer system 1200 50 logic 1226 during execution thereof by the computer system
within which a set of instructions, for causing the machine         1200 ; the main memory 1204 and the processing device
to perform any one or more of the methodologies discussed           1202 also constituting machine -accessible storage media .
herein , may be executed . In alternative embodiments , the            The machine - readable storage medium 1224 may also be
machine may be connected (e. g., networked ) to other used to store instructions 1226 implementing store address
machines in a LAN , an intranet, an extranet, or the Internet. 55 prediction for hybrid cores such as described according to
 The machine may operate in the capacity of a server or a embodiments of the disclosure . While the machine-acces
client device in a client- server network environment, or as a     sible storage medium 1128 is shown in an example embodi
peer machine in a peer -to - peer ( or distributed ) network       ment to be a single medium , the term “machine- accessible
environment. The machine may be a personal computer                storage medium ” should be taken to include a single
(PC ), a tablet PC , a set- top box (STB ), a Personal Digital 60 medium or multiple media ( e . g ., a centralized or distributed
Assistant (PDA ), a cellular telephone, a web appliance, a         database , and/ or associated caches and servers ) that store the
server, a network router, switch or bridge, or any machine one or more sets of instructions. The term “ machine - acces
capable of executing a set of instructions ( sequential or        s ible storage medium ” shall also be taken to include any
otherwise ) that specify actions to be taken by that machine . medium that is capable of storing , encoding or carrying a set
Further, while only a single machine is illustrated , the term 65 of instruction for execution by the machine and that cause
" machine ” shall also be taken to include any collection of        the machine to perform any one or more of the methodolo
machines that individually or jointly execute a set (or             gies of the present disclosure. The term “machine -accessible
                                                    US 10 ,313 , 108 B2
                              19                                                                20
storage medium ” shall accordingly be taken to include, but           compression employs 142 bits of the 256 -bit state data , and
not be limited to , solid - state memories, and optical and           round 60 of the compression employs 16 bits of the 256 -bit
magnetic media .                                                      state data .
   The following examples pertain to further embodiments .               In Example 7 , the subjectmatter of any of Examples of 1
Example 1 is a processing system includes a processor to 5 and 6 can further provide that the 1024 - bit input message
construct an inputmessage comprising a target value and a             further comprises a 256 -bit hash value recorded in a last
nonce and a hardware accelerator, communicatively coupled             block of a block chain recorded in a public ledger, a 256 - bit
to the processor, implementing a plurality of circuits to             Merkle root that is an initial hash value recorded in a first
perform stage - 1 secure hash algorithm (SHA ) hash and               block of the block chain , a 32-bit time stamp , and a plurality
stage - 2 SHA hash , wherein to perform the stage - 2 SHA " of padding bits .
hash , the hardware accelerator is to perform a plurality of   In Example 8 , the subjectmatter of Example 7 can further
rounds of compression on state data stored in a plurality of          provide that the SHA hash is a SHA -256 hash , wherein the
registers associated with a stage - 2 SHA hash circuit using an hardware accelerator is further to perform stage - O SHA hash
input value , wherein the input value comprises a hash value is on a first 512 bits of the 1024 -bits input message to generate
generated by a stage - 1 SHA hash circuit, and wherein each     a hash value which is used to initiate eight registers asso
register of the plurality of registers is to store a state that is ciated with the stage-1 SHA hash .
updated through the plurality of rounds of compression ,               In Example 9, the subject matter of Example 8 can further
calculate a plurality of speculative computation bits using a provide that the stage - 1 SHA hash is to receive a second 512
plurality of bits of the state data , and transmit the plurality 20 bits of the 1024-bit inputmessage as an input value to the
of speculative computation bits to the processor.                   stage - 1 SHA hash and to use the input value to the stage -1
   In Example 2 , the subjectmatter of Example 1 can further SHA hash to perform 64 rounds of compress on 256 -bit state
provide that the processor is to receive, from the hardware           data stored in the eight registers associated with stage - 1
accelerator, the plurality of speculative computation bits , SHA hash .
determine whether at least one bit of the plurality of specu - 25 Example 10 is an application specific integrated circuit
lative computation bits is non - zero , and responsive to deter - (ASIC ) comprising a plurality of registers and a plurality of
mining that at least one bit of the plurality of speculative      circuits to perform to perform stage - 1 secure hash algorithm
computation bits is non - zero , determine that the nonce is (SHA ) hash and stage - 2 SHA hash , wherein to perform the
invalid .                                                       stage- 2 SHA hash based on an inputmessage , the ASIC is
   In Example 3 , the subjectmatter of any of Examples 2 and 30 to perform a plurality of rounds of compression on state data
3 can further provide that the processor is to prior to               stored in a plurality of registers associated with a stage - 2
calculating the plurality of speculative computation bits,            SHA hash circuit using an input value , wherein the input
copy contents of the plurality of registers associated with the       value comprises a hash value generated by a stage - 1 SHA
stage - 2 SHA hash circuit to a second plurality of registers ,       hash circuit, and wherein each register of the plurality of
responsive to determining that all of the plurality of specu - 35 registers is to store a state that is updated through the
lative computation bits are zeros , copy contents of the plurality of rounds of compression , calculate a plurality of
second plurality of registers to the plurality of registers           speculative computation bits using a plurality of bits of the
associated with the stage - 2 SHA hash circuit , instruct the         state data , and transmit the plurality of speculative compu
hardware accelerator to perform additional rounds of the              tation bits to a processor communicatively coupled to the
compression using the stage - 2 SHA hash circuit to generate 40 ASIC .
a second hash value , receive , from the hardware accelerator ,         In Example 11 , the subject matter of Example 10 can
the second hash value, and compare the second hash value              further provide that the processor is to receive , from the
with the target value.                                                ASIC , the plurality of speculative computation bits , deter
   In Example 4 , the subjectmatter of Example 3 can further          mine whether at least one bit of the plurality of speculative
provide that the processor is to responsive to determining 45 computation bits is non -zero , and responsive to determining
that the second hash value is one of greater than or same as          that at least one bit of the plurality of speculative compu
the target value, determine that the nonce is invalid , and           tation bits is non -zero , determine that a nonce is invalid .
responsive to determining that the second hash value is                 In Example 12 , the subject matter of any of Examples 10
smaller than the target value, determine that the nonce is a          and 11 can further provide that the processor is further to
valid proof of identification of a Bitcoin coin .                  50 prior to calculating the plurality of speculative computation
   In Example 5 , the subject matter of Example 4 can further         bits , copy contents of the plurality of registers associated
provide that the processor is to responsive to determining            with the stage - 2 SHA hash circuit to a second plurality of
validity of the nonce , increment a value of the nonce to             registers , responsive to determining that all of the plurality
generate an updated inputmessage , and transmit the updated           of speculative computation bits are zeros, copy contents of
input message to the hardware accelerator to validate the 55 the second plurality of registers to the plurality of registers
incremented nonce .                                                   associated with the stage -2 SHA hash circuit , instruct the
  In Example 6 , the subjectmatter of Example 1 can further           ASIC to perform additional rounds of the compression using
provide that the input message comprises 1024 bits , the              the stage - 2 SHA hash circuit to generate a second hash
target value comprises 256 bit, and the nonce comprises 32            value , receive, from the ASIC , the second hash value, and
bits , wherein the plurality of rounds of compression com - 60 compare the second hash value with a target value.
prise fewer than 64 rounds , the plurality of registers com -           In Example 13 , the subject matter of Example 12 can
prise eight 32-bit registers to store 256 -bit state data , each      further provide that the processor is further to responsive to
32 -bit register storing a 32 -bit state, and wherein the plu -       determining that the second hash value is one of greater than
rality of speculative computation bits comprises two bits ,   or same as the target value , determine that the nonce is
and wherein round 57 of the compression employs 221 bits 65 invalid , and responsive to determining that the second hash
of the 256 -bit state data , round 58 of the compression      value is smaller than the target value , determine that the
employs 189 bits of the 256 -bit state data , round 59 of the nonce is a valid proof of identification of a Bitcoin coin .
                                                       US 10 ,313, 108 B2
                           21                                                                           22
   In Example 14 , the subject matter of Example 13 can                  of registers to the plurality of registers associated with the
further provide that responsive to determining validity of the           stage - 2 SHA hash circuit, instructing the hardware accel
nonce, increment a value of the nonce to generate an updated             erator to perform additional rounds of the compression using
inputmessage and transmit the updated inputmessage to the                the stage - 2 SHA hash circuit to generate a second hash
ASIC to validate the incremented nonce.                               5 value , receiving , from the hardware accelerator, the second
   In Example 15 , the subject matter of Example 14 can                  hash value, comparing the second hash value with the target
further provide that the input message comprises 1024 bits               value, responsive to determining that the second hash value
comprising a 256 -bit target value and a 32 -bit nonce,                  is one of greater than or same as the target value , determin
wherein the plurality of rounds of compression comprise                  ing that the nonce is invalid , and responsive to determining
fewer than 64 rounds, the plurality of registers comprise 10 that the second hash value is smaller than the target value ,
eight 32 -bit registers to store 256 -bit state data , each 32 -bit      determining that the nonce is a valid proof of identification
register storing a 32 -bit state , and wherein the plurality of          of a Bitcoin coin .
speculative computation bits comprises two bits , and                      In Example 21, an apparatus comprising: means for
wherein round 57 of the compression employs 221 bits of                  performing the method of any of Examples 18 to 20 .
the 256 -bit state data , round 58 of the compression employs 15            Example 22 is a machine-readable non - transitory medium
189 bits of the 256 -bit state data , round 59 of the compres            having stored thereon program code that, when executed ,
sion employs 142 bits of the 256 -bit state data , and round 60          perform operations comprising transmitting, by a processor,
of the compression employs 16 bits of the 256 -bit state data .          an inputmessage to a hardware accelerator, the inputmes
   In Example 16 , the subject matter of any of Examples 10              sage comprising a target value and a nonce , wherein the
and 15 can further provide that the ASIC is to receive, from 20 hardware accelerator implements a plurality of circuits to
the processor, a 1024 -bit input message comprising the                  perform stage - 1 secure hash algorithm (SHA ) hash and
256 - bit target value , the 32 -bit nonce , a 256 -bit hash value       stage - 2 SHA hash , instructing the hardware accelerator to
recorded in a last block of a block chain recorded in a public           perform plurality of rounds of compression on state data
ledger, a 256 -bit Merkle root that is an initial hash value              stored in plurality of registers associated with a stage - 2 SHA
recorded in a first block of the block chain , a 32 -bit time 25 hash circuit using an input value, wherein the input value
stamp, and a plurality of padding bits .                                 comprises a hash value generated by a stage- 1 SHA hash
  In Example 17 , the subject matter of Example 16 can                   circuit , and wherein each register of the plurality of registers
further provide that the SHA hash is a SHA - 256 hash ,                  is to store a state that is updated through the plurality of
wherein the ASIC is further to perform stage -0 SHA hash on              rounds of compression , instructing the hardware accelerator
a first 512 bits of the 1024 -bits inputmessage to generate a 30 to calculate a plurality of speculative computation bits using
hash value which is used to initiate eight registers associated  a plurality of bits of the state data , and receiving, from the
with the stage - 1 SHA hash , and the stage -1 SHA hash is to            hardware accelerator, the plurality of speculative computa
receive a second 512 bits of the 1024 - bit input message as            tion bits .
an input value to the stage- 1 SHA hash and to use the input                In Example 23 , the subject matter of Example 22 can
value to the stage- 1 SHA hash to perform 64 rounds of 35 further provide that the operations further comprise deter
compress on 256 -bit state data stored in the eight registers            mining whether at least one bit of the plurality of speculative
associated with stage - 1 SHA hash .                                     computation bits is non -zero , and responsive to determining
   Example 18 is a method comprising transmitting , by a                 that at least one bit of the plurality of speculative compu
processor, an input message to a hardware accelerator, the               tation bits is non - zero , determining that the nonce is invalid .
input message comprising a target value and a nonce , 40 In Example 24 , the subject matter of any of Examples 22
wherein the hardware accelerator implements a plurality of and 23 can further provide that the operations further
circuits to perform stage -1 secure hash algorithm (SHA )                comprise prior to calculating the plurality of speculative
hash and stage - 2 SHA hash , instructing the hardware accel-            computation bits , copying contents of the plurality of reg
erator to perform plurality of rounds of compression on state            isters associated with the stage - 2 SHA hash circuit to a
data stored in plurality of registers associated with a stage- 2 45 second plurality of registers , responsive to determining that
SHA hash circuit using an input value , wherein the input                all of the plurality of speculative computation bits are zeros ,
value comprises a hash value generated by a stage - 1 SHA                copying contents of the second plurality of registers to the
hash circuit, and wherein each register of the plurality of plurality of registers associated with the stage - 2 SHA hash
registers is to store a state that is updated through the        circuit, instructing the hardware accelerator to perform addi
plurality of rounds of compression , instructing the hardware 50 tional rounds of the compression using the stage - 2 SHA hash
accelerator to calculate a plurality of speculative computa -    circuit to generate a second hash value , receiving , from the
tion bits using a plurality of bits of the state data , and hardware accelerator, the second hash value , comparing the
receiving, from the hardware accelerator, the plurality of second hash value with the target value, responsive to
speculative computation bits .                                   determining that the second hash value is one of greater than
   In Example 19 , the subject matter of Example 18 can 55 or same as the target value, determining that the nonce is
further include determining whether at least one bit of the              invalid , and responsive to determining that the second hash
plurality of speculative computation bits is non -zero, and              value is smaller than the target value , determining that the
responsive to determining that at least one bit of the plurality         nonce is a valid proof of identification of a Bitcoin coin .
of speculative computation bits is non -zero , determining that            While the disclosure has been described with respect to a
the nonce is invalid .                                                60 limited number of embodiments , those skilled in the art will
   In Example 20 , the subjectmatter of any of Examples 18 appreciate numerous modifications and variations there
and 19 can further include prior to calculating the plurality   from . It is intended that the appended claims cover all such
of speculative computation bits , copying contents of the modifications and variations as fall within the true spirit and
plurality of registers associated with the stage- 2 SHA hash scope of this disclosure.
circuit to a second plurality of registers, responsive to 65 A design may go through various stages, from creation to
determining that all of the plurality of speculative compu - simulation to fabrication . Data representing a design may
tation bits are zeros, copying contents of the second plurality represent the design in a number of manners. First, as is
                                                    US 10 ,313 ,108 B2
                              23                                                                    24
useful in simulations, the hardware may be represented              Furthermore, use of the phrases 'to ,' ' capable of/to ,' and
using a hardware description language or another functional      or ' operable to ,' in one embodiment, refers to some appa
description language . Additionally , a circuit level model ratus, logic , hardware , and /or element designed in such a
with logic and /or transistor gates may be produced at some way to enable use of the apparatus, logic , hardware, and/or
 stages of the design process . Furthermore , most designs, at 5 element in a specified manner. Note as above that use of to ,
some stage , reach a level of data representing the physical capable to , or operable to , in one embodiment, refers to the
placement of various devices in the hardware model . In the wherelatent state of an apparatus, logic , hardware, and /or element,
case where conventional semiconductor fabrication tech operating         the apparatus, logic , hardware , and /or element is not
niques are used , the data representing the hardware model 10 an apparatus   but is designed in such a manner to enable use of
may be the data specifying the presence or absence of                            in a specified manner.
                                                                    A  value
various features on different mask layers for masks used to tion of a number , as used herein , includes any known representa
produce the integrated circuit. In any representation of the                         , a state, a logical state , or a binary logical
design , the data may be stored in any form of a machine             state . Often , the use of logic levels, logic values , or logical
                                                       values is also referred to as l ’ s and O ' s , which simply
readable medium . A memory or a magnetic or optical 15 represents binary logic states . For example, a 1 refers to a
storage such as a disc may be the machine readable medium           high logic level and O refers to a low logic level. In one
to store information transmitted via optical or electrical          embodiment, a storage cell, such as a transistor or flash cell ,
wave modulated or otherwise generated to transmit such              may be capable of holding a single logical value or multiple
information . When an electrical carrier wave indicating or logical values . However, other representations of values in
carrying the code or design is transmitted , to the extent that 20 computer systems have been used . For example the decimal
copying, buffering, or re -transmission of the electrical signal number ten may also be represented as a binary value of 910
is performed , a new copy is made . Thus, a communication            and a hexadecimal letter A . Therefore, a value includes any
provider or a network provider may store on a tangible ,             representation of information capable of being held in a
machine-readable medium , at least temporarily , an article ,        computer system .
such as information encoded into a carrier wave, embodying 25          Moreover, states may be represented by values or portions
techniques of embodiments of the present disclosure .                of values. As an example , a first value, such as a logical one ,
   A module as used herein refers to any combination of             may represent a default or initial state, while a second value,
hardware , software , and / or firmware. As an example , a          such as a logical zero , may represent a non - default state . In
module includes hardware , such as a micro - controller, asso -     addition , the terms reset and set, in one embodiment, refer
ciated with a non - transitory medium to store code adapted to 30 to a default and an updated value or state , respectively . For
be executed by the micro - controller. Therefore , reference to   example , a default value potentially includes a high logical
a module , in one embodiment, refers to the hardware, which          value , i.e . reset, while an updated value potentially includes
is specifically configured to recognize and /or execute the          a low logical value , i. e . set. Note that any combination of
code to be held on a non -transitory medium . Furthermore , in       values may be utilized to represent any number of states.
another embodiment, use of a module refers to the non - 35              The embodiments of methods, hardware , software , firm
transitory medium including the code , which is specifically         ware or code set forth above may be implemented via
adapted to be executed by the microcontroller to perform             instructions or code stored on a machine -accessible ,
predetermined operations . And as can be inferred , in yet          m achine readable , computer accessible , or computer read
another embodiment, the term module ( in this example) may          a ble medium which are executable by a processing element.
refer to the combination of the microcontroller and the 40 A non -transitory machine -accessible / readable medium
non - transitory medium . Often module boundaries that are           includes any mechanism that provides ( i.e ., stores and /or
illustrated as separate commonly vary and potentially over           transmits ) information in a form readable by a machine, such
lap . For example , a first and a second module may share            as a computer or electronic system . For example , a non
hardware, software, firmware, or a combination thereof,              transitory machine-accessible medium includes random -ac
while potentially retaining some independent hardware , 45 cess memory (RAM ), such as static RAM (SRAM ) or
software , or firmware . In one embodiment, use of the term          dynamic RAM (DRAM ) ; ROM ;magnetic or optical storage
logic includes hardware , such as transistors , registers , or      medium ; flash memory devices; electrical storage devices ;
other hardware , such as programmable logic devices .                optical storage devices; acoustical storage devices; other
   Use of the phrase " configured to ,' in one embodiment,           form of storage devices for holding information received
refers to arranging, putting together,manufacturing , offering 50 from transitory (propagated ) signals (e . g ., carrier waves,
to sell, importing and/ or designing an apparatus , hardware ,       infrared signals , digital signals ); etc ., which are to be dis
logic , or element to perform a designated or determined task .      tinguished from the non - transitory mediums that may
In this example, an apparatus or element thereof that is not         receive information there from .
operating is still 'configured to perform a designated task if          Instructions used to program logic to perform embodi
it is designed , coupled , and /or interconnected to perform 55 ments of the disclosure may be stored within a memory in
said designated task . As a purely illustrative example , a logic the system , such as DRAM , cache, flash memory , or other
gate may provide a 0 or a 1 during operation . But a logic gate   storage . Furthermore , the instructions can be distributed via
' configured to ' provide an enable signal to a clock does not       a network or by way of other computer readable media . Thus
include every potential logic gate that may provide a 1 or 0 . a machine -readable medium may include any mechanism
Instead , the logic gate is one coupled in some manner that 60 for storing or transmitting information in a form readable by
during operation the 1 or 0 output is to enable the clock . a machine (e . g ., a computer), but is not limited to , floppy
Note once again that use of the term ' configured to ' does not diskettes, optical disks, Compact Disc , Read -Only Memory
require operation , but instead focus on the latent state of an (CD -ROMs), and magneto - optical disks, Read -Only
apparatus, hardware , and / or element , where in the latent        Memory (ROMS), Random Access Memory (RAM ), Eras
state the apparatus, hardware , and /or element is designed to 65 able Programmable Read -Only Memory (EPROM ), Electri
perform a particular task when the apparatus , hardware ,         cally Erasable Programmable Read -Only Memory (EE
and /or element is operating .                                    PROM ), magnetic or optical cards, flash memory, or a
                                                      US 10 ,313,108 B2
                               25                                                                      26
tangible, machine -readable storage used in the transmission              responsive to determining that all of the plurality of
of information over the Internet via electrical, optical, acous              speculative computation bits are zeros, copy contents
tical or other forms of propagated signals ( e. g ., carrier                of the second plurality of registers to the plurality of
waves, infrared signals, digital signals, etc.). Accordingly,               registers associated with the stage- 2 SHA hash circuit ;
the computer - readable medium includes any type of tangible 5            instruct the hardware accelerator to perform additional
machine -readable medium suitable for storing or transmit                    rounds of the compression using the stage- 2 SHA hash
ting electronic instructions or information in a form readable               circuit to generate a second hash value ;
by a machine ( e . g ., a computer ).                                     receive , from the hardware accelerator, the second hash
  Reference throughout this specification to " one embodi                    value ; and
ment” or “ an embodiment” means that a particular feature, 10             compare the second hash value with the target value .
structure , or characteristic described in connection with the            4 . The processing system of claim 3, wherein the proces
embodiment is included in at least one embodiment of the               sor is further to :
present disclosure . Thus, the appearances of the phrases “ in            responsive to determining that the second hash value is
one embodiment” or “ in an embodiment” in various places                     one of greater than or same as the target value , deter
throughout this specification are not necessarily all referring 15          mine that the nonce is invalid ; and
to the same embodiment. Furthermore , the particular fea -                responsive to determining that the second hash value is
tures, structures, or characteristics may be combined in any       smaller than the target value , determine that the nonce
suitable manner in one or more embodiments .                       is a valid proof of identification of a Bitcoin coin .
   In the foregoing specification , a detailed description has 5 . The processing system of claim 4 , wherein the proces
been given with reference to specific exemplary embodi- 20 sor is further to :
ments. It will, however, be evident that various modifica                 responsive to determining validity of the nonce , incre
tions and changes may be made thereto without departing                     ment a value of the nonce to generate an updated input
from the broader spirit and scope of the disclosure as set                  message; and
forth in the appended claims. The specification and drawings              transmit the updated input message to the hardware
are , accordingly, to be regarded in an illustrative sense rather 25       accelerator to validate the incremented nonce .
than a restrictive sense . Furthermore , the foregoing use of           6 . The processing system of claim 1 , wherein the input
embodiment and other exemplarily language does not nec message comprises 1024 bits , the target value comprises 256
essarily refer to the same embodiment or the same example, bit, and the nonce comprises 32 bits, wherein the plurality of
butmay refer to different and distinct embodiments , as well         rounds of compression comprise fewer than 64 rounds, the
as potentially the same embodiment.                               30 plurality of registers comprise eight 32- bit registers to store
   What is claimed is:                                               256 -bit state data , each 32 -bit register storing a 32 -bit state ,
   1. A processing system comprising :                               and wherein the plurality of speculative computation bits
   a processor to construct an input message comprising a              comprises two bits , and wherein round 57 of the compres
      target value and a nonce ; and                           sion employs 221 bits of the 256 - bit state data , round 58 of
   a hardware accelerator, communicatively coupled to the 35 the compression employs 189 bits of the 256 -bit state data ,
      processor, implementing a plurality of circuits to per - round 59 of the compression employs 142 bits of the 256 -bit
      form stage - 1 secure hash algorithm (SHA ) hash and             state data , and round 60 of the compression employs 16 bits
      stage - 2 SHA hash based on the input message , wherein          of the 256 -bit state data .
      to perform the stage -2 SHA hash , the hardware accel             7 . The processing system of claim 6 , wherein the 1024 - bit
      erator is to :                                               40 input message further comprises a 256 -bit hash value
     perform a plurality of rounds of compression on state             recorded in a last block of a block chain recorded in a public
        data stored in a plurality of registers associated with        ledger, a 256 -bit Merkle root that is an initial hash value
        a stage - 2 SHA hash circuit using an input value ,            recorded in a first block of the block chain , a 32 -bit time
        wherein the input value comprises a hash value                 stamp, and a plurality of padding bits .
        generated by a stage - 1 SHA hash circuit , and 45 8 . The processing system of claim 7 , wherein the SHA
        wherein each register of the plurality of registers is hash is a SHA - 256 hash , and wherein the hardware accel
        to store a state that is updated through the plurality erator is further to perform stage - O SHA hash on a first 512
         of rounds of compression ;                                    bits of the 1024 -bits inputmessage to generate a hash value
      calculate a plurality of speculative computation bits            which is used to initiate eight registers associated with the
         using a plurality of bits of the state data ; and         50 stage - 1 SHA hash .
      transmit the plurality ofspeculative computation bits to            9. The processing system of claim 8 , wherein the stage - 1
         the processor                                                 SHA hash is to receive a second 512 bits of the 1024 - bit
   2 . The processing system of claim 1, wherein the proces             inputmessage as an input value to the stage - 1 SHA hash and
sor is to :                                                            to use the input value to the stage - 1 SHA hash to perform 64
   receive , from the hardware accelerator, the plurality of 55 rounds of compression on 256 -bit state data stored in the
      speculative computation bits ;                            eight registers associated with stage - 1 SHA hash .
   determine whether at least one bit of the plurality of                  10 . An application specific integrated circuit (ASIC )
      speculative computation bits is non -zero ; and                  comprising:
   responsive to determining that at least one bit of the                 a plurality of registers; and
     plurality of speculative computation bits is non -zero , 60          a plurality of circuits to perform stage - 1 secure hash
     determine that the nonce is invalid .                                    algorithm (SHA ) hash and stage - 2 SHA hash , wherein
   3. The processing system of claim 2 , wherein the proces                   to perform the stage - 2 SHA hash based on an input
sor is further to :                                                           message , the ASIC is to :
   prior to calculating the plurality of speculative computa                  perform a plurality of rounds of compression on state
      tion bits , copy contents of the plurality of registers 65                 data stored in a plurality of registers associated with
     associated with the stage - 2 SHA hash circuit to a                         a stage - 2 SHA hash circuit using an input value ,
      second plurality of registers ;                                           wherein the input value comprises a hash value
                                                          US 10 ,313 , 108 B2
                               27                                                                   28
         generated by a stage -1 SHA hash circuit , and a public ledger , a 256 -bit Merkle root that is an initial hash
         wherein each register of the plurality of registers is value recorded in a first block of the block chain , a 32 -bit
         to store a state that is updated through the plurality time stamp, and a plurality of padding bits .
         of rounds of compression ;                                    17 . The ASIC of claim 16 , wherein the SHA hash is a
      calculate a plurality of speculative computation bits hits 5 SHA -256 hash , and wherein the ASIC is further to perform
         using a plurality of bits of the state data ; and         stage - O SHA hash on a first 51 bits of the 1024 -bits input
                                                                   message     to generate a hash value which is used to initiate
       transmit the plurality of speculative computation bits to eight registers
         a processor communicatively coupled to the ASIC . stage - 1 SHA hash        associated with the stage - 1 SHAhash , and the
   11 . The ASIC of claim 10, wherein the processor is to :                                is to receive a second 512 bits of the
   receive , from the ASIC , the plurality of speculative com - 10 hash and to use the inputas value
                                                                   1024   -bit input message      an input value to the stage - 1 SHA
                                                                                                        to the stage - 1 SHA hash to
      putation bits ;                                              perform     64  rounds   of  compression   on 256 -bit state data
   determine whether at least one bit of the plurality of per       stored in the eight registers associated with stage - 1 SHA
         speculative computation bits is non - zero ; and
   responsive to determining that at least one bit of the hash .
      plurality of speculative computation bits is non -zero . 15 18 . A method comprising :
     determine that a nonce is invalid .                          transmitting , by a processor, an inputmessage to a hard
   12 . The ASIC of claim 11 ,wherein the processor is further       ware accelerator, the input message comprising a target
to :                                                                           value and a nonce, wherein the hardware accelerator
   prior to calculating the plurality of speculative computa                   implements a plurality of circuits to perform stage -1
      tion bits , copy contents of the plurality of registers 20               secure hash algorithm (SHA ) hash and stage -2 SHA
      associated with the stage - 2 SHA hash circuit to a                      hash ;
      second plurality of registers ;                                        instructing the hardware accelerator to perform a plurality
   responsive to determining that all of the plurality of                       of rounds of compression on state data stored in a
      speculative computation bits are zeros , copy contents                   plurality of registers associated with a stage - 2 SHA
      of the second plurality of registers to the plurality of 25              hash circuit using an input value , wherein the input
      registers associated with the stage -2 SHA hash circuit;                 value comprises a hash value generated by a stage- 1
   instruct the ASIC to perform additional rounds of the                       SHA hash circuit, and wherein each register of the
      compression using the stage - 2 SHA hash circuit to                       plurality of registers is to store a state that is updated
      generate a second hash value;                                             through the plurality of rounds of compression ;
       receive , from the ASIC , the second hash value ; and        30       instructing the hardware accelerator to calculate a plural
   compare the second hash value with a target value .                          ity of speculative computation bits using a plurality of
       13 . The ASIC of claim 12 , wherein the processor is further             bits of the state data ; and
to :                                                                         receiving, from the hardware accelerator, the plurality of
       responsive to determining that the second hash value is                 speculative computation bits .
         one of greater than or same as the target value, deter- 35          19 . The method of claim 18 , further comprising:
         mine that the nonce is invalid ; and                                determining whether at least one bit of the plurality of
   responsive to determining that the second hash value is                     speculative computation bits is non - zero ; and
          smaller than the target value, determine that the nonce            responsive to determining that at least one bit of the
           is a valid proof of identification of a Bitcoin coin .              plurality of speculative computation bits is non -zero ,
       14 . The ASIC of claim 13 , wherein the processor is further 40          determining that the nonce is invalid .
to :                                                                         20 . The method of claim 19 , further comprising:
       responsive to determining validity of the nonce, incre                prior to calculating the plurality of speculative computa
     ment a value of the nonce to generate an updated input                     tion bits , copying contents of the plurality of registers
     message ; and                                                              associated with the stage - 2 SHA hash circuit to a
   transmit the updated input message to the ASIC to vali - 45                 second plurality of registers ;
      date the incremented nonce .                                           responsive to determining that all of the plurality of
       15 . The ASIC of claim 10 , wherein the input message                   speculative computation bits are zeros, copying con
comprises 1024 bits comprising a 256 -bit target value and a                   tents of the second plurality of registers to the plurality
32 -bit nonce , wherein the plurality of rounds of compression                 of registers associated with the stage - 2 SHA hash
comprise fewer than 64 rounds, the plurality of registers 50                   circuit ;
comprise eight 32 -bit registers to store 256 -bit state data ,              instructing the hardware accelerator to perform additional
each 32 - bit register storing a 32 -bit state, and wherein the                 rounds of the compression using the stage-2 SHA hash
plurality of speculative computation bits comprises two bits ,                 circuit to generate a second hash value ;
and wherein round 57 of the compression employs 221 bits                     receiving, from the hardware accelerator, the second hash
of the 256 -bit state data , round 58 of the compression 55                    value;
employs 189 bits of the 256 - bit state data , round 59 of the               comparing the second hash value with the target value;
compression employs 142 bits of the 256 -bit state data , and                responsive to determining that the second hash value is
round 60 of the compression employs 16 bits of the 256 -bit                    one of greater than or same as the target value, deter
state data .                                                                   mining that the nonce is invalid ; and
   16 . The ASIC of claim 15 , whereinin the ASIC isis toto receive
                                         the ASIC           receive , 6060   responsive to determining that the second hash value is
                                                                                smaller than the target value , determining that the
from the processor, the 1024 -bit inputmessage comprising
the 256 -bit target value , the 32 -bit nonce , a 256 -bit hash                nonce is a valid proof of identification ofa Bitcoin coin .
value recorded in a last block of a block chain recorded in