Back to ZeliDesk MainZeliDesk 메인으로 돌아가기返回 ZeliDesk 首頁

Find the AI Model Format for Your PC내 PC에 맞는 AI 모델 포맷 찾기尋找適合你電腦的 AI 模型格式

ZeliDesk sLM supports various hardware environments. Select your specs below to get optimal format and inference engine recommendations. All major formats and quantization techniques are documented for academic reference.ZeliDesk sLM은 다양한 하드웨어 환경을 지원합니다. 아래에서 사양을 선택하면 최적의 포맷과 추론 엔진을 추천받을 수 있습니다. 학계·연구 참고 자료로도 활용 가능하도록 현존하는 모든 주요 포맷과 양자화 기법을 수록했습니다.ZeliDesk sLM 支援各種硬體環境。選擇您的規格後,即可獲得最佳格式與推論引擎建議。所有主要格式與量化技術均已收錄,可作為學術參考。

16+
Formats수록 포맷收錄格式
5
Categories카테고리分類
10+
Paper Refs논문 레퍼런스論文參考

Configure Your PC Specs내 PC 사양 설정設定你的電腦規格

Select each option to enable analysis.각 항목을 선택하면 분석 버튼이 활성화됩니다.選擇各項目後,分析按鈕將啟用。

Graphics Card (GPU)그래픽 카드 (GPU)顯示卡 (GPU)
No GPU (Integrated)GPU 없음 (내장 그래픽)無 GPU(內建顯示)Intel UHD / AMD APU etc.Intel UHD / AMD APU 등Intel UHD / AMD APU 等
NVIDIA Low-endNVIDIA 구형/저사양NVIDIA 入門級GTX 1050~1660, RTX 2060
NVIDIA Mid-rangeNVIDIA 중급NVIDIA 中階RTX 3060~3070, RTX 4060
NVIDIA High-endNVIDIA 고급NVIDIA 高階RTX 3080+, RTX 4070~4090
NVIDIA Data CenterNVIDIA 데이터센터NVIDIA 資料中心A100, H100, Blackwell
AMD RadeonRX 6000~7000 시리즈
Intel ArcArc A750~A770
Apple SiliconM1~M4 시리즈
VRAM (Graphics Memory)VRAM (그래픽 메모리)VRAM(顯示記憶體)
None / Unknown없음 / 모름無 / 不確定
4 GB or less4 GB 이하4 GB 以下
6 GB
8 GB
12 GB
16 GB
24 GB
48 GB+48 GB 이상48 GB 以上
Unified Memory (Apple)통합 메모리 (Apple)統一記憶體(Apple)RAM = VRAM sharedRAM = VRAM 공유RAM = VRAM 共享
System RAM시스템 RAM系統 RAM
8 GB or less8 GB 이하8 GB 以下
16 GB
32 GB
64 GB
128 GB+128 GB 이상128 GB 以上
Usage Purpose사용 목적使用目的
Speed First속도 우선速度優先Fastest response가장 빠른 응답最快回應
🔧 Compatibility호환성 우선相容性優先Works everywhere어떤 환경에서도 동작任何環境皆可運作
🧠 Quality First품질 우선品質優先Model accuracy & quality모델 정확도·응답 품질模型準確度·回應品質
📦 Save Space용량 절약節省空間Low disk space디스크 공간 부족磁碟空間不足
🔬 Research연구/학습研究/學習Fine-tuning · Experiments파인튜닝·실험·논문微調·實驗·論文
🌐 Server Deploy서버 배포伺服器部署Production serving프로덕션 서빙生產環境服務

Analysis Results분석 결과分析結果

Model Format & Quantization Encyclopedia모델 포맷 · 양자화 기법 백과사전模型格式·量化技術百科

All major formats and quantization techniques organized by category. Includes paper references.현존하는 모든 주요 포맷과 양자화 기법을 카테고리별로 정리했습니다. 논문 레퍼런스 포함.所有主要格式與量化技術按類別整理。包含論文參考。

Distribution & Storage Formats배포 · 저장 포맷部署·儲存格式 Container for models모델을 담는 그릇模型的容器
GGUF
llama.cpp · Universal local deployment standardllama.cpp · 범용 로컬 배포 표준llama.cpp · 通用本地部署標準
The de facto standard of the local AI ecosystem. Embeds metadata, tokenizer, and architecture info in a single file. Runs on CPU alone, with broad GPU acceleration support for NVIDIA CUDA / AMD Vulkan·ROCm / Apple Metal. Most flexible RAM+VRAM offloading.현재 로컬 AI 생태계의 사실상 표준. 메타데이터·토크나이저·아키텍처 정보를 단일 파일에 내장. CPU만으로도 구동 가능하며, NVIDIA CUDA / AMD Vulkan·ROCm / Apple Metal 등 폭넓은 GPU 가속을 지원. RAM+VRAM 혼용(Offloading)이 가장 자유로움.本地 AI 生態系的事實標準。將元資料、分詞器與架構資訊內嵌於單一檔案中。僅用 CPU 即可運行,並廣泛支援 NVIDIA CUDA / AMD Vulkan·ROCm / Apple Metal 等 GPU 加速。RAM+VRAM 混合使用最為靈活。
📄 GGML → GGUF 전환: Georgi Gerganov, llama.cpp (2023~)
Compatibility호환성相容性★★★★★
GPU SpeedGPU 속도GPU 速度★★★☆☆
Safetensors
Hugging Face · Research/sharing standardHugging Face · 연구/공유 표준Hugging Face · 研究/共享標準
The de facto standard storage format for AI research. Solves security issues of Pickle-based formats. Supports 3-tier offloading (VRAM→RAM→Disk) with device_map="auto". AWQ/GPTQ/BitsAndBytes quantized models also use this format.AI 연구 커뮤니티의 사실상 표준 저장 포맷. Pickle 기반 포맷(.bin)의 보안 문제를 해결. device_map="auto"로 VRAM→RAM→디스크 3단계 오프로딩 가능. AWQ/GPTQ/BitsAndBytes 양자화 모델도 이 포맷을 사용.AI 研究社群的事實標準儲存格式。解決了 Pickle 格式的安全問題。透過 device_map="auto" 支援 VRAM→RAM→磁碟三層卸載。AWQ/GPTQ/BitsAndBytes 量化模型也使用此格式。
📄 Hugging Face Safetensors (2022)
Compatibility호환성相容性★★★★☆
GPU SpeedGPU 속도GPU 速度★★★☆☆
ONNX
Microsoft · Cross-platform industry standardMicrosoft · 크로스 플랫폼 산업 표준Microsoft · 跨平台產業標準
ZeliDesk's default embedded format. Supports diverse hardware via CPU/CUDA/DirectML/OpenVINO Execution Providers. Strong for enterprise and edge device deployment. Provides INT4 quantized models (Phi-3 etc.) as built-in.ZeliDesk 기본 탑재 포맷. CPU/CUDA/DirectML/OpenVINO 등 다양한 Execution Provider로 폭넓은 하드웨어 지원. 엔터프라이즈와 엣지 디바이스 배포에 강점. INT4 양자화 모델(Phi-3 등)을 내장형으로 제공.ZeliDesk 預設內建格式。透過 CPU/CUDA/DirectML/OpenVINO 等執行提供者支援多種硬體。適合企業與邊緣裝置部署。提供 INT4 量化模型(Phi-3 等)作為內建。
📄 ONNX: Open Neural Network Exchange (2017~, Microsoft + Meta)
Compatibility호환성相容性★★★★☆
GPU SpeedGPU 속도GPU 速度★★★☆☆
GGML (레거시)
GGUF predecessor · Historical referenceGGUF의 전신 · 역사적 참고GGUF 前身 · 歷史參考
Format used before GGUF. Had limitations of relying on external files for metadata, replaced by GGUF. Found in pre-2023 model distributions. Not recommended for new use.GGUF 이전에 사용되던 포맷. 메타데이터를 외부 파일에 의존하는 한계가 있어 GGUF로 대체됨. 2023년 이전의 구형 모델 배포본에서 볼 수 있음. 현재 신규 사용은 권장하지 않음.GGUF 之前使用的格式。因依賴外部檔案存放元資料的限制而被 GGUF 取代。可見於 2023 年前的舊版模型分發。目前不建議新使用。
📄 Georgi Gerganov, ggml (2022)
GPU Quantization — ProductionGPU 양자화 — 프로덕션GPU 量化 — 生產環境 NVIDIA GPU OptimizedNVIDIA GPU 최적화NVIDIA GPU 最佳化
AWQ
Activation-Aware Weight Quantization
The de facto standard for production INT4 quantization as of 2026. Protects 'salient' weights to minimize quantization loss. Officially supported by major serving engines including vLLM, SGLang, TensorRT-LLM. Highly VRAM-dependent.2026년 현재 프로덕션 INT4 양자화의 사실상 표준. "중요한(salient)" 가중치를 보호하여 양자화 손실을 최소화. vLLM, SGLang, TensorRT-LLM 등 주요 서빙 엔진에서 공식 지원. VRAM 의존도 높음.截至 2026 年生產環境 INT4 量化的事實標準。保護「顯著」權重以最小化量化損失。獲 vLLM、SGLang、TensorRT-LLM 等主要服務引擎官方支援。高度依賴 VRAM。
📄 Lin et al., "AWQ: Activation-aware Weight Quantization for LLM Compression and Acceleration" (MLSys 2024)
Quality Retention품질 보존品質保留★★★★★
GPU SpeedGPU 속도GPU 速度★★★★☆
GPTQ
Post-Training Quantization for GPT
One of the oldest LLM-specific quantization techniques. Layer-wise quantization based on OBQ (Optimal Brain Quantization). Was the GPU quantization standard before AWQ, still widely used. High-speed inference via Marlin kernels.가장 오래된 LLM 전용 양자화 기법 중 하나. OBQ(Optimal Brain Quantization)를 기반으로 한 레이어별 양자화. AWQ 등장 전까지 GPU 양자화의 표준이었으며, 여전히 광범위하게 사용. Marlin 커널을 통해 고속 추론 가능.最早的 LLM 專用量化技術之一。基於 OBQ(最佳大腦量化)的逐層量化。在 AWQ 出現前是 GPU 量化標準,至今仍廣泛使用。可透過 Marlin 核心實現高速推論。
📄 Frantar et al., "GPTQ: Accurate Post-Training Quantization for Generative Pre-trained Transformers" (ICLR 2023)
Quality Retention품질 보존品質保留★★★★☆
GPU SpeedGPU 속도GPU 速度★★★★☆
EXL2
ExLlamaV2 · Consumer GPU optimizedExLlamaV2 · 소비자 GPU 최적화ExLlamaV2 · 消費級 GPU 最佳化
Currently the fastest inference speed on consumer NVIDIA GPUs. Fine-grained quantization in bpw (bits-per-weight) units (2.5~8.0bpw) for precise VRAM fitting. Model must fit entirely in VRAM. No RAM offloading.소비자 NVIDIA GPU에서 현재 가장 빠른 추론 속도. bpw(bits-per-weight) 단위의 세밀한 양자화(2.5~8.0bpw)로 VRAM에 딱 맞게 조절 가능. 모델이 반드시 VRAM 안에 들어가야 함. RAM 오프로딩 불가.目前消費級 NVIDIA GPU 上最快的推論速度。以 bpw(每權重位元)為單位的精細量化(2.5~8.0bpw),可精確匹配 VRAM。模型必須完全載入 VRAM。不支援 RAM 卸載。
📄 turboderp, ExLlamaV2 (2023~)
Quality Retention품질 보존品質保留★★★★☆
GPU SpeedGPU 속도GPU 速度★★★★★
Marlin
High-performance GPU inference kernel고성능 GPU 추론 커널高效能 GPU 推論核心
Highly optimized CUDA kernels for GPTQ/AWQ quantized models. Core backend of vLLM. Specialized for FP16×INT4 mixed operations, matching EXL2 throughput in simple decoding.GPTQ/AWQ 양자화 모델을 위한 고도로 최적화된 CUDA 커널. vLLM의 핵심 백엔드로 사용. FP16×INT4 혼합 연산에 특화되어, 단순 디코딩에서는 EXL2에 필적하는 처리량을 보여줌.為 GPTQ/AWQ 量化模型高度最佳化的 CUDA 核心。vLLM 的核心後端。專精於 FP16×INT4 混合運算,在簡單解碼中可匹敵 EXL2 的吞吐量。
📄 IST Austria, "Marlin: Mixed-Precision Auto-Regressive Parallel INference" (2024)
Throughput처리량吞吐量★★★★★
FP8
8-bit Floating Point · Next-gen standard8-bit Floating Point · 차세대 표준8 位浮點數 · 下一代標準
Hardware native support on latest NVIDIA GPUs (Hopper H100/Blackwell). Maintains near-BF16 quality while using half the memory. The 'practical sweet spot' — higher quality than INT4, faster than FP16. Applicable to both training and inference.Hopper(H100)/Blackwell 등 최신 NVIDIA GPU의 하드웨어 네이티브 지원. BF16에 근접한 품질을 유지하면서 메모리 절반 사용. INT4보다 품질이 높고, FP16보다 빠른 "실용적 최적점". 학습과 추론 모두에 적용.最新 NVIDIA GPU(Hopper H100/Blackwell)的硬體原生支援。維持接近 BF16 的品質,同時使用一半記憶體。「實用最佳點」——品質高於 INT4,速度快於 FP16。適用於訓練與推論。
📄 NVIDIA, "FP8 Formats for Deep Learning" (2022); Micikevicius et al.
Quality Retention품질 보존品質保留★★★★★
Hardware Req.하드웨어 요구硬體需求Hopper+ 필수
NVFP4
NVIDIA 4-bit Floating Point
Latest NVIDIA-exclusive 4-bit floating point format introduced with Blackwell architecture. Higher precision than INT4 while pushing memory bandwidth to the limit. Still in early stages, full support expected via TensorRT-LLM.Blackwell 아키텍처에서 도입된 최신 NVIDIA 전용 4비트 부동소수점 포맷. INT4보다 정밀도가 높으면서도 메모리 대역폭을 극한까지 줄임. 아직 초기 단계이며, 향후 TensorRT-LLM을 통해 본격 지원 예정.Blackwell 架構引入的最新 NVIDIA 專用 4 位浮點格式。精度高於 INT4,同時將記憶體頻寬推至極限。仍處於早期階段,預計透過 TensorRT-LLM 全面支援。
📄 NVIDIA Blackwell Architecture Whitepaper (2024~2025)
Research Quantization연구용 양자화 기법研究用量化技術 Paper-based · Extreme compression논문 기반 · 극한 압축論文基礎·極限壓縮
BitsAndBytes (NF4/INT8)
Core of QLoRA fine-tuningQLoRA 파인튜닝의 핵심QLoRA 微調的核心
The 'one-line' quantization standard of the Hugging Face ecosystem. Apply NF4 quantization with just load_in_4bit=True. The core tool for QLoRA fine-tuning. Most widely used in research and prototyping. Focused on convenience over inference speed.Hugging Face 생태계의 "원라인" 양자화 표준. load_in_4bit=True 한 줄로 NF4 양자화 적용. QLoRA 파인튜닝의 핵심 도구. 연구·프로토타이핑에서 가장 많이 사용. 추론 속도보다 편의성에 초점.Hugging Face 生態系的「一行」量化標準。僅需 load_in_4bit=True 即可套用 NF4 量化。QLoRA 微調的核心工具。在研究與原型開發中使用最廣。著重便利性而非推論速度。
📄 Dettmers et al., "QLoRA: Efficient Finetuning of Quantized Language Models" (NeurIPS 2023)
Convenience편의성便利性★★★★★
Inference Speed추론 속도推論速度★★☆☆☆
HQQ
Half-Quadratic Quantization
Innovative technique enabling quantization without calibration data. Maintains high quality at ultra-low bits (2-3 bits) via half-quadratic optimization. Faster quantization than GPTQ. A promising next-gen technique in research.캘리브레이션 데이터 없이 양자화 가능한 혁신적 기법. 반이차(Half-Quadratic) 최적화로 극저비트(2~3비트)에서도 높은 품질 유지. GPTQ보다 빠른 양자화 속도. 연구에서 주목받는 차세대 기법.無需校準資料即可量化的創新技術。透過半二次最佳化在超低位元(2-3 位元)維持高品質。量化速度比 GPTQ 更快。研究中備受矚目的下一代技術。
📄 Badri & Shaji, "HQQ: Half-Quadratic Quantization of Large Language Models" (2023)
Ultra-low bit quality극저비트 품질超低位元品質★★★★☆
AQLM
Additive Quantization of Language Models
Multi-codebook based quantization. Represents weights by summing multiple codebooks. Maintains vastly superior perplexity over GPTQ/AWQ at 2-bit level. Very long quantization time but top-tier quality.다중 코드북 기반 양자화. 여러 개의 코드북을 합산(Additive)하여 가중치를 표현. 2비트 수준에서 GPTQ/AWQ보다 월등한 perplexity 유지. 양자화 시간이 매우 길지만 품질은 최상급.基於多碼本的量化。透過多個碼本相加來表示權重。在 2 位元水準下維持遠優於 GPTQ/AWQ 的困惑度。量化時間很長但品質頂級。
📄 Egiazarian et al., "AQLM: Extreme Compression of Language Models via Additive Quantization" (ICML 2024)
2-bit quality2비트 품질2 位元品質★★★★★
QuIP#
Quantization with Incoherence Processing
Theoretically sophisticated quantization combining lattice codebooks and incoherence processing. Achieves state-of-the-art perplexity at 2-bit. High academic influence, inspiring many follow-up works.격자 코드북(Lattice Codebook)비일관성 처리(Incoherence Processing)를 결합한 이론적으로 정교한 양자화. 2비트에서 state-of-the-art perplexity를 달성. 학술적 영향력이 크며, 후속 연구에 큰 영감을 줌.理論精密的量化技術,結合格子碼本非相干處理。在 2 位元達到最先進的困惑度。學術影響力高,啟發許多後續研究。
📄 Tseng et al., "QuIP#: Even Better LLM Quantization with Hadamard Incoherence and Lattice Codebooks" (ICML 2024)
SpQR
Sparse-Quantized Representation
A sparse-quantization hybrid that keeps outlier weights at high precision while quantizing the rest at ultra-low bits. Based on the insight that less than 1% of outlier weights are critical for quality.이상치(Outlier) 가중치를 고정밀도로 유지하고 나머지를 극저비트로 양자화하는 희소-양자화 하이브리드. 전체 모델의 1% 미만인 이상치 가중치가 품질에 결정적이라는 통찰에 기반.以高精度保留離群權重,其餘以超低位元量化的稀疏-量化混合方法。基於不到 1% 的離群權重對品質至關重要的洞見。
📄 Dettmers et al., "SpQR: A Sparse-Quantized Representation for Near-Lossless LLM Weight Compression" (ICLR 2024)
SqueezeLLM
Dense-and-Sparse Quantization
Combines sensitivity-based non-uniform quantization with sparse decomposition. Excellent perplexity retention at 3-bit. The Dense-and-Sparse concept influenced subsequent production frameworks.감도 기반(Sensitivity-based) 비균일 양자화 + 희소 분해(Sparse Decomposition)를 결합. 3비트 양자화에서 뛰어난 perplexity 유지. Dense-and-Sparse 개념이 후속 프로덕션 프레임워크에 영향.結合基於敏感度的非均勻量化與稀疏分解。在 3 位元量化下保持優異的困惑度。Dense-and-Sparse 概念影響了後續的生產框架。
📄 Kim et al., "SqueezeLLM: Dense-and-Sparse Quantization" (ICML 2024)
MXFP (Microscaling)
Next-gen block floating point차세대 블록 부동소수점下一代區塊浮點數
Next-generation numerical format co-proposed by Microsoft/AMD/Intel/NVIDIA. Shares scaling factors at block level for high precision at ultra-low bits. Expected to receive hardware native support in the future.Microsoft/AMD/Intel/NVIDIA 등이 공동 제안한 차세대 수치 형식. 블록 단위로 스케일링 팩터를 공유하여 극저비트에서도 높은 정밀도를 유지. 향후 하드웨어 네이티브 지원이 예상되는 포맷.Microsoft/AMD/Intel/NVIDIA 共同提出的下一代數值格式。在區塊層級共享縮放因子,在超低位元下維持高精度。預計未來獲得硬體原生支援。
📄 OCP Microscaling Formats Specification (2023); Rouhani et al.
Hardware-Specific Formats하드웨어 전용 포맷硬體專用格式 Chipset-specific optimization특정 칩셋 최적화特定晶片最佳化
TensorRT-LLM
NVIDIA Official · Extreme optimizationNVIDIA 공식 · 극한 최적화NVIDIA 官方 · 極致最佳化
NVIDIA's official inference engine. Compiles models specifically for the user's exact GPU architecture, achieving theoretical peak performance. Features In-flight Batching, FP8/INT4 quantization, KV Cache optimization.NVIDIA 공식 추론 엔진. 사용자의 정확한 GPU 아키텍처에 맞춰 모델을 컴파일(빌드)하기 때문에 이론적 최고 성능을 달성. In-flight Batching, FP8/INT4 양자화, KV Cache 최적화 등 엔터프라이즈 기능 탑재.NVIDIA 官方推論引擎。針對使用者確切的 GPU 架構編譯模型,達到理論上的最高效能。具備 In-flight Batching、FP8/INT4 量化、KV Cache 最佳化等企業功能。
📄 NVIDIA TensorRT-LLM (2023~)
GPU SpeedGPU 속도GPU 速度★★★★★+
Setup Difficulty세팅 난이도設定難度★★★★★
MLX
Apple Silicon · Unified Memory optimizedApple Silicon · 통합 메모리 최적화Apple Silicon · 統一記憶體最佳化
Apple's framework for M-series chips. Leverages Unified Memory for inference without CPU/GPU data copying. Better battery efficiency and speed than GGUF on Mac. PyTorch-like API.Apple이 만든 M시리즈 칩 전용 프레임워크. 통합 메모리(Unified Memory)를 활용하여 CPU/GPU 데이터 복사 없이 추론. Mac 환경에서 GGUF보다 더 나은 배터리 효율과 속도. PyTorch와 유사한 API.Apple 為 M 系列晶片打造的框架。運用統一記憶體進行推論,無需 CPU/GPU 資料複製。在 Mac 上比 GGUF 有更好的電池效率與速度。類似 PyTorch 的 API。
📄 Apple Machine Learning Research, "MLX" (2023~)
Mac PerformanceMac 성능Mac 效能★★★★★
OpenVINO
Intel · CPU/GPU/NPU optimizedIntel · CPU/GPU/NPU 최적화Intel · CPU/GPU/NPU 最佳化
Intel hardware (Core Ultra NPU, Arc GPU, Xeon CPU) specific optimization toolkit. Auto-optimizes ONNX/PyTorch models for Intel hardware. Particularly strong for edge device and IoT deployment.Intel 하드웨어(Core Ultra NPU, Arc GPU, Xeon CPU) 전용 최적화 툴킷. ONNX/PyTorch 모델을 Intel 하드웨어에 맞게 자동 최적화. 엣지 디바이스·IoT 배포에 특히 강점.Intel 硬體(Core Ultra NPU、Arc GPU、Xeon CPU)專用最佳化工具包。自動將 ONNX/PyTorch 模型針對 Intel 硬體最佳化。特別適合邊緣裝置與 IoT 部署
📄 Intel OpenVINO Toolkit (2018~)
CoreML
Apple · iOS/macOS nativeApple · iOS/macOS 네이티브Apple · iOS/macOS 原生
Native model format for iPhone/iPad/Mac's ANE (Apple Neural Engine). Used when embedding AI models directly into iOS apps. Unlike MLX, inference-only and optimized for app distribution.iPhone/iPad/Mac의 ANE(Apple Neural Engine)에서 구동되는 네이티브 모델 포맷. iOS 앱에 AI 모델을 직접 탑재할 때 사용. MLX와 달리 추론 전용이며, 앱 배포에 최적화.iPhone/iPad/Mac ANE(Apple 神經引擎)原生模型格式。用於將 AI 模型直接嵌入 iOS 應用程式。與 MLX 不同,僅供推論且針對應用分發最佳化。
📄 Apple CoreML (2017~)
Specialized Inference Engines특수 추론 엔진 · 프레임워크特殊推論引擎·框架 Unique runtime environments독자적 실행 환경獨特執行環境
BitNet / bitnet.cpp
1.58-bit Ternary models1.58-bit 삼진(Ternary) 모델1.58 位元三元模型
Revolutionary architecture using only {-1, 0, 1} weight values. Trained in ternary from the start, not post-quantized. Even 100B parameter models can run fast on CPU alone. Dramatic energy reduction. At the forefront of AI efficiency research.가중치가 {-1, 0, 1} 세 값만 사용하는 혁명적 아키텍처. 양자화가 아닌 처음부터 삼진법으로 학습. 100B 파라미터 모델도 CPU만으로 빠르게 구동 가능. 에너지 소비 극적 감소. AI 효율성 연구의 최전선.革命性架構,權重僅使用 {-1, 0, 1} 三個值從一開始就以三元法訓練,非後量化。即使 1000 億參數模型也能僅用 CPU 快速運行。大幅降低能源消耗。AI 效率研究的最前線。
📄 Ma et al., "The Era of 1-bit LLMs: All Large Language Models are in 1.58 Bits" (Microsoft Research, 2024)
CPU EfficiencyCPU 효율CPU 效率★★★★★
Model Variety모델 다양성模型多樣性★★☆☆☆
MLC-LLM
Machine Learning Compilation
Compiler-based universal inference engine. Auto-optimizes and deploys one model across NVIDIA/AMD/Apple/WebGPU/Android. Based on TVM framework. Core technology behind WebLLM (running LLMs in browsers).컴파일러 기반 범용 추론 엔진. 하나의 모델을 NVIDIA/AMD/Apple/WebGPU/Android 등 다양한 플랫폼에 자동 최적화하여 배포. TVM 프레임워크 기반. 브라우저에서 LLM을 돌리는 WebLLM의 핵심 기술.基於編譯器的通用推論引擎。自動將一個模型最佳化並部署至 NVIDIA/AMD/Apple/WebGPU/Android 等多種平台。基於 TVM 框架。在瀏覽器中運行 LLM 的 WebLLM 核心技術。
📄 CMU Catalyst Group, "MLC-LLM" (2023~); Apache TVM
PowerInfer
Sparsity-based hybrid inference희소성 기반 하이브리드 추론基於稀疏性的混合推論
Predicts neuron activation patterns, placing frequently used 'hot neurons' on GPU and the rest on CPU. Much more efficient RAM+VRAM mixing than conventional layer-wise offloading. Advantageous for running large models on consumer GPUs.뉴런 활성화 패턴을 예측하여, 자주 쓰이는 "핫 뉴런"은 GPU에, 나머지는 CPU에 배치. 일반적인 레이어별 오프로딩보다 훨씬 효율적인 RAM+VRAM 혼용. 소비자 GPU에서 큰 모델을 돌릴 때 유리.預測神經元激活模式,將常用的「熱神經元」放在 GPU 上,其餘放在 CPU 上。比傳統逐層卸載更高效的 RAM+VRAM 混合使用。有利於在消費級 GPU 上運行大型模型。
📄 Song et al., "PowerInfer: Fast Large Language Model Serving with a Consumer-grade GPU" (OSDI 2024)
FlexGen
Offloading-specialized inference engineOffloading 전문 추론 엔진卸載專用推論引擎
Maximally utilizes GPU/CPU/Disk 3-tier hierarchical memory to run ultra-large models (175B) on a single GPU. Optimized for batch processing. Suited for bulk inference rather than real-time conversation.GPU/CPU/디스크의 3단계 계층적 메모리를 최대한 활용하여 단일 GPU에서 초대형 모델(175B)을 구동하는 오프로딩 전문 엔진. 배치 처리에 최적화. 실시간 대화보다는 대량 추론에 적합.最大化利用 GPU/CPU/磁碟三層階層式記憶體,在單一 GPU 上運行超大模型(175B)。針對批次處理最佳化。適合大量推論而非即時對話。
📄 Sheng et al., "FlexGen: High-Throughput Generative Inference of Large Language Models with a Single GPU" (ICML 2023)
vLLM
Server serving standard engine서버 서빙 표준 엔진伺服器服務標準引擎
Server inference engine maximizing throughput via PagedAttention for KV Cache memory management. Supports most quantization formats including AWQ/GPTQ/FP8/Marlin. The de facto standard for production LLM serving.PagedAttention으로 KV Cache 메모리를 페이지 단위로 관리하여 처리량을 극대화하는 서버용 추론 엔진. AWQ/GPTQ/FP8/Marlin 등 대부분의 양자화 포맷 지원. 프로덕션 LLM 서빙의 사실상 표준.透過 PagedAttention 進行 KV Cache 記憶體管理,最大化吞吐量的伺服器推論引擎。支援 AWQ/GPTQ/FP8/Marlin 等大多數量化格式。生產環境 LLM 服務的事實標準
📄 Kwon et al., "Efficient Memory Management for Large Language Model Serving with PagedAttention" (SOSP 2023)
Server Throughput서버 처리량伺服器吞吐量★★★★★
Ollama
One-click local AI execution원클릭 로컬 AI 실행一鍵本地 AI 執行
Download and run GGUF models with a single command like Docker: ollama run llama3. Uses llama.cpp internally. The lowest barrier tool for anyone to experience local AI instantly.GGUF 모델을 Docker처럼 간편하게 ollama run llama3 한 줄로 다운로드·실행하는 도구. 내부적으로 llama.cpp를 사용. 비개발자도 로컬 AI를 즉시 체험할 수 있는 진입 장벽 최저 도구.像 Docker 一樣用一行指令下載並執行 GGUF 模型:ollama run llama3。內部使用 llama.cpp。讓任何人都能立即體驗本地 AI 的最低門檻工具
📄 Ollama (2023~); Jeffrey Morgan et al.

All Formats at a Glance전체 포맷 한눈에 비교所有格式一覽比較

16+ formats and quantization techniques organized at a glance.16개 이상의 포맷·양자화 기법을 일목요연하게 정리합니다.16 種以上格式與量化技術一目了然。

Format포맷/기법格式/技術 Engine추론 엔진推論引擎 GPU Req.GPU 필수需要 GPU RAM+VRAM MixRAM+VRAM 혼용RAM+VRAM 混用 NVIDIA AMD Apple Intel Recommended추천 환경推薦環境
📦 Distribution · Storage Formats📦 배포 · 저장 포맷📦 分發·儲存格式
GGUFllama.cpp✅ Best✅ 최고✅ 最佳✅ CUDA✅ Vulkan✅ Metal⚠️All PC levels구형~고급 PC 모두所有 PC 等級
SafetensorsHF Transformers✅ 3-tier✅ 3단계✅ 三層✅ CUDA✅ ROCm✅ MPS⚠️Researchers · Python연구자 · Python 생태계研究者 · Python
ONNXONNX Runtime⚠️ Limited⚠️ 제한⚠️ 有限✅ CUDA✅ DML✅ CoreML✅ OpenVINOWindows · EnterpriseWindows · 엔터프라이즈Windows · 企業
🚀 GPU 양자화 — 프로덕션
AWQvLLM / AutoAWQ⚠️프로덕션 서빙 (표준)
GPTQAutoGPTQ / Marlin⚠️GPU 서버 · RTX 3060+
EXL2ExLlamaV2소비자 NVIDIA (8GB↑)
MarlinvLLM 커널고처리량 서빙
FP8TRT-LLM / vLLM✅ H100+데이터센터 (Hopper+)
NVFP4TRT-LLM✅ Blackwell차세대 DC (Blackwell)
🔬 Research Quantization🔬 연구용 양자화🔬 研究用量化
BitsAndBytesHF Transformers⚠️⚠️QLoRA fine-tuningQLoRA 파인튜닝QLoRA 微調
HQQHF / NativeHF / 자체HF / 自己⚠️Calibration-free quantization캘리브레이션 없는 양자화無校準量化
AQLMHF Transformers⚠️Ultra-low bit (2bit) research극저비트(2bit) 연구超低位元(2bit)研究
QuIP#Custom code자체 코드自己程式Ultra-low bit academic research극저비트 학술 연구超低位元學術研究
SpQRCustom code자체 코드自己程式Outlier weight research이상치 가중치 연구離群權重研究
SqueezeLLMCustom code자체 코드自己程式Dense-Sparse researchDense-Sparse 연구Dense-Sparse 研究
🖥️ 하드웨어 전용
TensorRT-LLMNVIDIA TRT✅ 전용NVIDIA DC · 고급 사용자
MLXApple MLXApple통합메모리✅ 네이티브Apple Silicon (M1~M4)
OpenVINOIntel OV⚠️✅ 네이티브Intel CPU/NPU/Arc
CoreMLApple CoreMLApple통합메모리✅ ANEiOS/macOS 앱 배포
🔧 특수 엔진
BitNetbitnet.cppN/A⚠️⚠️⚠️⚠️CPU 초효율 추론
PowerInfer자체 엔진⚠️✅ 예측 기반저사양 GPU 대형 모델
vLLMPagedAttention✅ ROCm프로덕션 서빙 표준
MLC-LLMTVM 컴파일러⚠️✅ Vulkan✅ Metal⚠️크로스플랫폼 · WebGPU