The line and the shelf
Every performance threshold in US computer export controls since 1992, converted to FP64 and set against what one person could buy. India is on the chart twice, 38 years apart.
What the chart is and is not
The two lines are not measured in the same unit, and a single axis through both is a device. The 1990s regimes counted composite theoretical performance on general-purpose 64-bit arithmetic. The 2006 rules replaced that with a weighted count of 64-bit floating-point operations. The 2022 rules count multiply-accumulates at whatever precision the chip runs fastest, which for a modern accelerator is 8-bit integer, not 64-bit floating point. A reader who catches that before the author admits it stops trusting the rest, so it is stated here first.
The one continuous thing underneath is FP64: the arithmetic a Cray did, a transputer did, and a laptop core still does. Every threshold on the chart has been converted to FP64 as follows. Millions of theoretical operations per second (MTOPS) count as MFLOPS one for one, because the CTP formula's word-length factor for a 64-bit processor is 1.0.1 Weighted TeraFLOPS (WT) are divided by the 0.9 weight the rule assigns to vector processors; the 0.3 weight for other processors would put the line three times higher.2 The 2022 chip threshold, TPP 4,800, is divided by 64 bits to give 75 TFLOPS, the FP64 rate that would trip it if FP64 were the chip's fastest precision. No shipping chip is built that way, so the dashed line is an upper bound on where the rule sits in FP64 terms, not a measurement of anything.3
Threshold values come from the regulations and the Federal Register notices that changed them, not from summaries. The consumer-shelf points come from Jack Dongarra's LINPACK report where a measured or theoretical FP64 figure exists, from manufacturer datasheets for graphics cards, and from this strand's own benchmark for the current laptop. Each point carries its source class in the data file.
1987
In 1986 India asked to buy a Cray X-MP/24 for the National Centre for Medium Range Weather Forecasting. The stated use was monsoon prediction. Washington declined the two-processor machine and offered the single-processor X-MP/14 instead, on conditions: civilian use only, no use for nuclear-weapons research, and US personnel with access to the installation. India objected to the personnel condition, then withdrew the X-MP/24 request days before Rajiv Gandhi's visit to Washington and took the X-MP/14. The agreement was signed on 9 October 1987 by Ambassador John Gunther Dean and Foreign Secretary K. P. S. Menon. India became the first country outside the Western alliance and Aramco to be cleared for the machine.4
An X-MP CPU ran a 9.5 nanosecond clock and 64-bit words.5 Dongarra's table lists the single-CPU X-MP/14se at 53 MFLOPS on LINPACK 100, 184 MFLOPS on LINPACK 1000, and a theoretical peak of 210 MFLOPS, all FP64.6 The refused X-MP/24 had two such CPUs. That is the point marked at 420 MFLOPS on the chart, above every consumer box of its decade by three orders of magnitude.
What India built instead
C-DAC was set up in 1988. Its first machine, PARAM 8000, was delivered as a 64-node system in August 1991. Each node was an Inmos T800-class transputer, a 32-bit microprocessor with an on-chip 64-bit floating-point unit and four serial links, wired through a reconfigurable interconnect.7
The published numbers for that hardware are small. Dongarra lists a single 20 MHz T800 under Fortran at 0.37 MFLOPS on LINPACK 100, FP64.6 Inmos's own technical note gives 1.5 MFLOPS for the T800-20 on Livermore Loop 7, but that figure is single-length, 32-bit, and is labelled as such wherever it appears in this strand.8 Sixty-four T800s at 0.37 MFLOPS each is 23.7 MFLOPS if every node solves its own system and none of them talk, which is a ceiling, not a benchmark. C-DAC's own figures for a 256-node machine were a theoretical 1 GFLOPS and a sustained 100 to 200 MFLOPS; those come to us through Kahaner's 1996 survey and are recorded here as reported, not as measured.9
Two claims travel with PARAM 8000 in almost every account. One is that it was 28 times the power of the Cray India was refused, for the same $10 million. The other is that a prototype came second only to a US machine at a supercomputing show in Zurich in 1990. Both originate with C-DAC or its supporters. The first appears in V. Rajaraman's 1999 book; the second in a 1998 magazine profile, and Wikipedia's own editors note that the event cannot be attested beyond that article.10 They are interesting as claims. They do not appear in any table on this site.
The documented outcome is the one that matters for the argument. The machine sold. C-DAC's accounts and contemporary press list installations in Germany, the United Kingdom and Russia, including a system for ICAD Moscow in 1991, at a list price around $350,000.11 A denial produced a domestic alternative at a fraction of the price, and the alternative was then exported to countries that ran the control regime. That fact sits uncomfortably with both sides of the current argument about accelerators, and it is rarely to hand when that argument is made.
The line, 1992 to 2023
Composite theoretical performance, CTP, became the yardstick in the early 1990s. Commerce defined a high-performance computer as 195 MTOPS in 1992 and 1,500 MTOPS in 1994.12 In January 1996 the United States sorted the world into four computer tiers. India, with China, Russia and Pakistan, was Tier 3: no licence below 2,000 MTOPS, a licence exception to civilian end-users up to 7,000, and a licence above that.13 After the May 1998 nuclear tests, exports above 2,000 MTOPS to listed Indian entities carried a presumption of denial.14
Then the line moved, repeatedly. 6,500 MTOPS in February 2000. 12,500 in August 2000. 28,000 in February 2001. 85,000 in May 2001. 190,000 in March 2002.15 Five increases in 25 months, each justified in the Federal Register by the availability of the hardware below it. In 2006 MTOPS was retired for Adjusted Peak Performance, a weighted count of 64-bit floating-point operations, and the control level for a digital computer became 0.75 Weighted TeraFLOPS.16 That level went to 1.5 WT in 2011, 3.0 in 2012, 8.0 in 2014, 12.5 in 2016, 16 in 2017, 29 in 2018 and 70 in 2023.17
Against that line, the shelf. A 1996 Pentium Pro had a theoretical FP64 peak of 200 MFLOPS, over the 1992 definition of a supercomputer within four years. A 2001 Athlon at 1.4 GHz peaked at 2,800 MFLOPS, over the 1996 Tier 3 line within five. A 2007 Core 2 Quad, 38.4 GFLOPS, cleared the 12,500 and 28,000 MTOPS lines of 2000 and 2001. A 2011 Core i7 at 108.8 GFLOPS cleared the 85,000 line after ten years. A 2013 GeForce GTX Titan, sold to anyone with $999, offered 1.3 TFLOPS of double precision and cleared both the 190,000 MTOPS line of 2002 and the 0.75 WT line of 2006. A 2019 Radeon VII at 3.46 TFLOPS cleared the 1.5 and 3.0 WT lines.18
Then the pattern breaks, and the break is the finding. No consumer part has cleared the 8.0 WT line of 2014 in FP64. Consumer graphics silicon since 2019 runs double precision at a small fraction of single, between a sixteenth and a sixty-fourth depending on the vendor; the 2022 RTX 4090 offers 1.29 TFLOPS FP64, less than the 2019 Radeon. The ten performance cores of the M4 Pro this page was benchmarked on reach 0.37 TFLOPS on Livermore Loop 7, and about 0.72 TFLOPS if each core's four 128-bit SIMD pipes are counted at full FMA rate, which Apple does not publish and this piece treats as an estimate. In the unit the old regime measured, the consumer shelf stopped climbing around 2019 while the line kept rising to 78 TFLOPS. Measured in FP64, the interval between drawing a line and crossing it did not keep shrinking. It went to infinity.
The line, 2022 to now
The regime that followed does not measure FP64. The October 2022 rule created ECCN 3A090 for chips with a total processing performance of 4,800 or more, together with an interconnect of 600 GB/s or more; TPP is two times multiply-accumulates per second times the bit length of the operation, taken at the chip's fastest precision.3 The November 2023 revision dropped the interconnect test and added performance density.19 The RTX 4090 had been on shelves since 12 October 2022. At 8-bit precision its TPP is above 4,800. Under the 2023 revision it was a controlled item with 1.29 TFLOPS of double precision, and Nvidia shipped a reduced 4090D for China, as widely reported in December 2023. Measured in the regime's own unit, the consumer shelf was above the line thirteen months before the line was drawn where it could catch it.
India's second appearance is at the other end of the chart. The Framework for Artificial Intelligence Diffusion, published 15 January 2025, extended the licence requirement for 3A090.a chips to every destination and sorted countries into three tiers. India was in the second, with a per-country ceiling on controlled compute through 2027.20 The framework was rescinded in May 2025 before its compliance date, and the tiers with it; the chip thresholds stayed.21 In January 2026 the review policy for chips below TPP 21,000 destined for China moved from presumption of denial to case-by-case.22 The argument made in 1987 about a weather computer, that a general-purpose machine of a certain speed is a weapon input and its buyer must accept conditions, was made again in 2025 about a rack of accelerators and the same country, for four months.
Numbers
The table below is generated from this strand's benchmark repository. Measured cells were run on this laptop with the compiler and flags shown; published cells resolve to a citation; the one reconstructed cell is labelled. FP32 figures carry the label inside the cell. Full data, harness and sources: param-fp64.
| machine | mode | thermal | LL7 | LL7 (fp32) | LINPACK 100 | LINPACK 1000 | peak | sustained |
|---|---|---|---|---|---|---|---|---|
| t800-30 | single | — | 2.25 (fp32) ᵖ | — | — | — | — | |
| t800-20 | single | — | 1.50 (fp32) ᵖ | 0.37 ᵖ | — | — | — | |
| t414-20 | single | — | 0.09 (fp32) ᵖ | — | — | — | — | |
| vax11-780-fpa | single | — | 0.54 (fp32) ᵖ | — | — | — | — | |
| cray-xmp14se | single | — | — | 53 ᵖ | 184 ᵖ | 210 ᵖ | — | |
| cray-xmp416-1cpu | single | — | — | 121 ᵖ | 218 ᵖ | 235 ᵖ | — | |
| param8000-64 | throughput | — | — | 24 ʳ | — | — | — | |
| param8000-256 | throughput | — | — | — | — | 1,000 ᵖ | 150 ᵖ | |
| m4pro-1p | single | cold | 40,936 | 76,788 (fp32) | 15,128 | 11,931 | — | — |
| m4pro-1p | single | sustained | 40,955 | 76,794 (fp32) | 15,143 | 11,960 | — | — |
| m4pro-10p | throughput | sustained | 365,360 | 658,504 (fp32) | 135,641 | 86,803 | — | — |
MFLOPS, FP64 unless labelled. ᵖ published · ʳ reconstructed · unmarked measured. Measured cells: Apple clang version 21.0.0 (clang-2100.1.1.101), flags -O2 -std=c99 -fno-fast-math, median of 2 runs for sustained cells; throughput = 10 independent processes, MFLOPS summed.
A word on the ratios these invite. On LINPACK 100, FP64, one M4 Pro performance core against one T800-20 is a factor of about forty thousand; against one X-MP/14se CPU, about three hundred. On Livermore Loop 7 at 32-bit, the precision Inmos published, the factor over a T800-20 is about fifty thousand. The X-MP figure is the more instructive one: the class of machine that was a diplomatic question in 1987 is outrun three hundred to one by a single core of a laptop bought at retail, in the same arithmetic.
Sources
- Category 4 Technical Note on CTP, 15 CFR 774 Supp. 1 (pre-2006 text): CTP = R × L, L = 1/3 + WL/96, so L = 1.0 at a 64-bit word length. Values in MTOPS are therefore read as MFLOPS for a 64-bit floating-point processor. As implemented by 61 FR 2099 (25 Jan 1996).
- Technical Note on "Adjusted Peak Performance", 15 CFR 774 Supp. 1, Category 4, as in force 1 Sept 2026 (eCFR): "APP" = Σ Wi × Ri, R the peak 64-bit floating-point rate, W = 0.9 for vector processors and 0.3 otherwise.
- 87 FR 62186 (13 Oct 2022, effective 7 Oct 2022), ECCN 3A090.a: "aggregate bidirectional transfer rate … of 600 Gbyte/s or more" and "bit length per operation multiplied by processing performance measured in TOPS … of 4800 or more." The TPP technical note in 88 FR 73458: "the 'TPP' threshold of 4800 can be met with 600 tera integer operations … at 8 bits or 300 tera FLOPS … at 16 bits."
- UPI, "U.S. signs first sale of supercomputer to non-Western nation", 9 October 1987. upi.com.
- Cray Research, The CRAY X-MP Series of Computer Systems, 1985 (MP-0102): X-MP/1 "9.5 nsec clock cycle time", "64-bit words". Computer History Museum, PDF. The brochure states no MFLOPS figure; 2 results per clock at 105.3 MHz gives the 210 MFLOPS peak Dongarra prints.
- J. J. Dongarra, Performance of Various Computers Using Standard Linear Equations Software, CS-89-85, edition of 15 June 2014, Table 1. netlib.org. Rows: "Cray X-MP/14se (10 ns) cf77 3.0 53 184 210"; "Inmos T800 (20 MHz) Fortran 3L -:o0 .37". Full precision, 64-bit.
- C-DAC's founding in 1988 and the node description are from Wikipedia, PARAM, citing D. K. Kahaner, "Parallel computing in India", IEEE Parallel & Distributed Technology 4(3), 1996, doi:10.1109/88.532134, as cited in Wikipedia, PARAM (fetched 21 Sept 2026). The paper itself was not read for this piece; the figures are relayed through that citation.
- Inmos Technical Note 6, IMS T800 Architecture, January 1988: "The IMS T800-30 achieves a speed of 2.25 Mflops on this benchmark; for comparison the IMS T800-20 achieves 1.5 Mflops, the T414-20 achieves 0.09 Mflops and a VAX 11/780 (with fpa) achieves 0.54 Mflops." The accompanying code uses the single-length load instructions. transputer.net.
- Kahaner 1996, as in note 7: "A 256-node machine had a theoretical performance of 1 GFLOPS, however in practice had a sustained performance of 100–200 MFLOPS."
- V. Rajaraman, Super computers, Universities Press (India), 1999, p. 75; Outlook Business, "God, Man And Machine", 1 July 1998; both as cited in Wikipedia, PARAM, including the editorial note on the Zurich claim.
- The Hindu Business Line, 26 Feb 2001; C-DAC, "From PARAM 8000 to PARAM 10000" (ICAD Moscow); Washington Post archive on the $350,000 price and 14 buyers; all as cited in Wikipedia, PARAM. C-DAC is the source for its own export list.
- Congressional Research Service, RL31175, High Performance Computers and Export Control Policy: Issues for Congress: "In 1992, the U.S. Commerce Department defined an HPC as 195 MTOPS … revised in 1994 (1,500 MTOPS)". Secondary; the 1992 and 1994 notices predate the online Federal Register.
- 61 FR 2099 (25 Jan 1996): Tier 3 "General License G-DEST for computers less than or equal to 2,000 MTOPS"; "G-CTP for computers greater than 2,000 MTOPS but less than or equal to 7,000 MTOPS"; validated licence above 7,000.
- 63 FR 64322 (19 Nov 1998), India and Pakistan Sanctions: "presumption of denial for all applications for exports and reexports of computers having a CTP greater than 2,000 MTOPS destined to Indian and Pakistani entities determined to be involved in nuclear…".
- 64 FR 42009 (3 Aug 1999; 2,000 → 6,500 military, 7,000 → 12,300 civilian, effective Feb 2000); 65 FR 12919 (10 Mar 2000; 12,500 and 20,000, effective 14 Aug 2000); 65 FR 60852 (13 Oct 2000; 28,000, civil–military distinction removed, effective 26 Feb 2001); 66 FR 5443 (19 Jan 2001; 85,000, effective 19 May 2001); 67 FR 10608 (8 Mar 2002; 190,000, effective 3 Mar 2002).
- 71 FR 20876 (24 Apr 2006), Implementation of New Formula for Calculating Computer Performance: Adjusted Peak Performance (APP): 4A003.b at 0.75 WT.
- 76 FR 36986 (24 Jun 2011; 0.75 → 1.5 WT); 77 FR 39354 (2 Jul 2012; 3.0); 79 FR 45288 (4 Aug 2014; 8.0); 81 FR 64656 (20 Sept 2016; 12.5); 82 FR 38764 (15 Aug 2017; 16, effective 25 Sept 2017); 83 FR 53742 (24 Oct 2018; 29); 88 FR 12108 (24 Feb 2023; 70). Current text confirmed in eCFR as of 1 Sept 2026.
- Dongarra 2014, Table 1: "Gateway 2000 G6-200 PentiumPro … 62 … 200"; "AMD Athlon Thunderbird 1.4GHz … 704 … 2800"; "Intel Core 2 Q6600 … (4 core, 2.4 GHz) … 13130 … 38400". Core i7-2600K: 4 cores × 8 FP64 flops per cycle (256-bit AVX add and multiply) × 3.4 GHz = 108.8 GFLOPS, computed from the architecture, not measured. GTX Titan: Nvidia newsroom, 19 Feb 2013, "4.5 teraflops of single-precision and 1.3 teraflops of double-precision". Radeon VII: AMD figure of 3.46 TFLOPS FP64 as reported by Techgage, Feb 2019. RTX 4090: 82.58 TFLOPS FP32 at a 1:64 FP64 ratio = 1.29 TFLOPS, per the Ada Lovelace architecture documentation as reported by TechPowerUp. M4 Pro: this strand's benchmark, see the table above.
- 88 FR 73458 (25 Oct 2023, effective 17 Nov 2023), ECCN 3A090.a: "a 'total processing performance' of 4800 or more, or … 1600 or more and a 'performance density' of 5.92 or more"; 3A090.b added.
- 90 FR 4544 (15 Jan 2025, effective 13 Jan 2025), Framework for Artificial Intelligence Diffusion. India's tier and the per-country allocation are stated in the rule's country lists and §740.29.
- BIS announced the rescission of the AI Diffusion rule on 13 May 2025, ahead of its 15 May compliance date; reported by Kirkland & Ellis, 27 May 2025, among others. Secondary.
- 91 FR 1684 (15 Jan 2026), Revision to License Review Policy for Advanced Computing Commodities: case-by-case review for chips with TPP below 21,000 and DRAM bandwidth below 6,500 GB/s to China and Macau.