Skip to main content
QUICK REVIEW

[논문 리뷰] A Unified Cryptoprocessor for Lattice-based Signature and Key-exchange

Aikata Aikata, Ahmet Can Mert|arXiv (Cornell University)|2022. 10. 13.
Cryptography and Data Security인용 수 7
한 줄 요약

이 논문은 라티스 기반 양자후계 암호화를 위한 통합적이고 컴act하며 프로그래머블한 암호처리기 아키텍처를 제안한다. 이 아키텍처는 디지털 서명을 위한 CRYSTALS-Dilithium과 키 캡슐화 메커니즘을 위한 Saber를 통합한다. 공통된 알고리즘적 구성요소—특히 통합된 NTT 기반 다항식 승수와 재구성 가능한 Keccak 코어—를 활용하여 최소한의 면적 오버헤드로 높은 성능을 달성하였으며, FPGA에서 200 MHz로 동작하고, 두 방식 모두에서 핵심 연산의 실행 시간이 100 μs 이내임을 입증하였다.

ABSTRACT

We propose design methodologies for building a compact, unified and programmable cryptoprocessor architecture that computes post-quantum key agreement and digital signature. Synergies in the two types of cryptographic primitives are used to make the cryptoprocessor compact. As a case study, the cryptoprocessor architecture has been optimized targeting the signature scheme 'CRYSTALS-Dilithium' and the key encapsulation mechanism (KEM) 'Saber', both finalists in the NIST's post-quantum cryptography standardization project. The programmable cryptoprocessor executes key generations, encapsulations, decapsulations, signature generations, and signature verifications for all the security levels of Dilithium and Saber. On a Xilinx Ultrascale+ FPGA, the proposed cryptoprocessor consumes 18,406 LUTs, 9,323 FFs, 4 DSPs, and 24 BRAMs. It achieves 200 MHz clock frequency and finishes CCA-secure key-generation/encapsulation/decapsulation operations for LightSaber in 29.6/40.4/58.3$μ$s; for Saber in 54.9/69.7/94.9$μ$s; and for FireSaber in 87.6/108.0/139.4$μ$s, respectively. It finishes key-generation/sign/verify operations for Dilithium-2 in 70.9/151.6/75.2$μ$s; for Dilithium-3 in 114.7/237/127.6$μ$s; and for Dilithium-5 in 194.2/342.1/228.9$μ$s, respectively, for the best-case scenario. On UMC 65nm library for ASIC the latency is improved by a factor of two due to a 2x increase in clock frequency.

연구 동기 및 목표

  • 양자후계 키 캡슐화(Saber)와 디지털 서명(Dilithium) 방식을 모두 지원하는 컴act하고 통합적이며 프로그래머블한 암호처리기를 설계하는 것.
  • NTT 우호적(Dilithium)과 NTT 비우호적(Saber) 라티스 기반 방식 간의 알고리즘적 및 아키텍처적 상호보완성을 활용하여 하드웨어 오버헤드를 최소화하는 것.
  • 자원 제약 환경(예: TLS 기반 시스템)에서의 실용적 구현을 위해 고성능과 낮은 면적 소비를 달성하는 것.
  • 모듈러 인스트럭션 세트 아키텍처를 통해 향후 라티스 기반 방식인 CRYSTALS-Kyber와의 확장성을 보장하는 것.
  • 공통 구성요소에 대한 마스킹 전략을 통합하여 사이드 채널에 강건한 구현 기반을 마련하는 것.

제안 방법

  • 공통 하드웨어 유닛을 사용하여 Dilithium 및 Saber 연산을 모두 지원하는 통합 인스트럭션 세트 아키텍처를 설계하는 것.
  • NTT 우호적(Dilithium)과 최적화된 NTT 비우호적(Saber) 산술 연산을 모두 지원하는 단일 NTT 기반 다항식 승수를 구현하는 것. 이는 특수 소수 모듈러스를 활용한다.
  • SHA3 및 SHAKE 연산을 둘 다 사용하는 데에 필요한 실시간 입력 전처리 및 출력 후처리를 효율적으로 처리하기 위해 재구성 가능한 Keccak 코어 워핑을 개발하는 것.
  • 파이ipel라인 데이터플로우 및 버퍼링 전략을 통해 중복 읽기/쓰기 사이클을 줄임으로써 메모리 액세스 패턴을 최적화하는 것.
  • 양 방식 간에 공유되는 다항식 승수 및 모듈로 감소 연산을 위한 이중 모드 산술 유닛을 통합하는 것.
  • 런타임 재구성 및 다양한 보안 수준을 지원하기 위해 하드웨어-소프트웨어 공동 설계 접근법을 사용하는 것.
Figure 1: Distribution of a coefficient after a multiplication between a secret and a public polynomial. The secret and public coefficients are in the range $[-\frac{\mu}{2},\frac{\mu}{2}]$ and $[0,q-1]$ , respectively.
Figure 1: Distribution of a coefficient after a multiplication between a secret and a public polynomial. The secret and public coefficients are in the range $[-\frac{\mu}{2},\frac{\mu}{2}]$ and $[0,q-1]$ , respectively.

실험 결과

연구 질문

  • RQ1NTT 우호적과 비우호적 라티스 방식 간의 알고리즘적 차이에도 불구하고, 단일 암호처리기가 라티스 기반 키 캡슐화(Saber)와 디지털 서명(Dilithium)을 효율적이고 컴팩트하게 지원할 수 있는가?
  • RQ2NTT 우호적 및 비우호적 라티스 방식 간의 아키텍처적 상호보완성은 무엇이며, 이를 통해 면적 오버헤드를 최소화할 수 있는가?
  • RQ3통합된 하드웨어 설계는 키 생성 및 서명 검증과 같은 핵심 연산에서 고속도를 유지하면서 자원 소비를 어떻게 줄일 수 있는가?
  • RQ4NTT 및 Keccak 코어와 같은 공통 빌딩 블록은 성능을 저하시키지 않고 다양한 양자후계 방식 간에 얼마나 효과적으로 재사용될 수 있는가?
  • RQ5제안된 암호처리기는 CRYSTALS-Kyber와 같은 추가적인 라티스 기반 방식을 최소한의 수정으로 지원할 수 있는가?

주요 결과

  • 암호처리기는 Xilinx Ultrascale+ FPGA에서 200 MHz 클럭 주파수를 달성하였으며, 18,406 LUT, 9,323 FF, 4 DSP, 24 BRAM을 소비하였다.
  • LightSaber의 경우, 키 생성, 캡슐화, 디캡슐화가 각각 29.6, 40.4, 58.3 μs 내에 완료되었다.
  • Saber의 경우, 동일한 연산은 각각 54.9, 69.7, 94.9 μs 소요되었으며, Dilithium-3 서명 및 검증은 최적의 경우 237 및 127.6 μs 내에 완료되었다.
  • 65nm ASIC에서 클럭 주파수가 2배로 증가함에 따라 성능이 두 배로 향상되었으며, Saber 연산의 추정 지연 시간은 각각 27.5, 34.9, 47.5 μs였다.
  • 특수 소수 모듈러스를 활용하여 Saber의 NTT 계산을 최적화함으로써, Dilithium 전용 설계 대비 면적 오버헤드를 감소시켰다.
  • 최적화된 Keccak 워핑은 불필요한 메모리 액세스를 줄여 해시 기반 연산의 처리량을 향상시키고 전력 소비를 감소시켰다.
Figure 2: The high-level design of the cryptoprocessor with the Saber, Dilithium and common modules are in green, blue and red, respectively.
Figure 2: The high-level design of the cryptoprocessor with the Saber, Dilithium and common modules are in green, blue and red, respectively.

더 나은 연구,지금 바로 시작하세요

논문 읽기부터 검토까지, 연구 시간을 획기적으로 줄여보세요.

카드 등록 없음 · 무료 플랜 제공

이 리뷰는 AI가 만들고, 인간 에디터가 검토했습니다.