Hardware & Computational Tools Directory
Showing 193 - 240 of 20,000 verified tools.
Stable Diffusion XL 6.6B Base (INT8 SmoothQuant Precision) on NVIDIA A100 80GB PCIe VRAM & Throughput Calculator
Exact VRAM memory allocation, dynamic KV-cache requirements, and tensor parallelism slicing for Stable Diffusion XL 6.6B Base quantized in INT8 SmoothQuant Precision deployed on NVIDIA A100 80GB PCIe.
Stable Diffusion XL 6.6B Base (AWQ 4-Bit Activation-Aware) on NVIDIA RTX 4090 24GB GDDR6X VRAM & Throughput Calculator
Exact VRAM memory allocation, dynamic KV-cache requirements, and tensor parallelism slicing for Stable Diffusion XL 6.6B Base quantized in AWQ 4-Bit Activation-Aware deployed on NVIDIA RTX 4090 24GB GDDR6X.
Stable Diffusion XL 6.6B Base (GPTQ 4-Bit Second-Order) on NVIDIA L40S 48GB Ada Lovelace VRAM & Throughput Calculator
Exact VRAM memory allocation, dynamic KV-cache requirements, and tensor parallelism slicing for Stable Diffusion XL 6.6B Base quantized in GPTQ 4-Bit Second-Order deployed on NVIDIA L40S 48GB Ada Lovelace.
Stable Diffusion XL 6.6B Base (GGUF Q4_K_M Medium Quant) on AMD Instinct MI300X 192GB VRAM & Throughput Calculator
Exact VRAM memory allocation, dynamic KV-cache requirements, and tensor parallelism slicing for Stable Diffusion XL 6.6B Base quantized in GGUF Q4_K_M Medium Quant deployed on AMD Instinct MI300X 192GB.
CogVideoX-5B Video Synthesis (FP16 Uncompressed Native) on NVIDIA H100 80GB SXM5 VRAM & Throughput Calculator
Exact VRAM memory allocation, dynamic KV-cache requirements, and tensor parallelism slicing for CogVideoX-5B Video Synthesis quantized in FP16 Uncompressed Native deployed on NVIDIA H100 80GB SXM5.
CogVideoX-5B Video Synthesis (BF16 Bfloat16 Mixed Precision) on NVIDIA H200 141GB HBM3e VRAM & Throughput Calculator
Exact VRAM memory allocation, dynamic KV-cache requirements, and tensor parallelism slicing for CogVideoX-5B Video Synthesis quantized in BF16 Bfloat16 Mixed Precision deployed on NVIDIA H200 141GB HBM3e.
CogVideoX-5B Video Synthesis (FP8 Scaled Native Hopper) on NVIDIA B200 192GB Blackwell VRAM & Throughput Calculator
Exact VRAM memory allocation, dynamic KV-cache requirements, and tensor parallelism slicing for CogVideoX-5B Video Synthesis quantized in FP8 Scaled Native Hopper deployed on NVIDIA B200 192GB Blackwell.
CogVideoX-5B Video Synthesis (INT8 SmoothQuant Precision) on NVIDIA A100 80GB PCIe VRAM & Throughput Calculator
Exact VRAM memory allocation, dynamic KV-cache requirements, and tensor parallelism slicing for CogVideoX-5B Video Synthesis quantized in INT8 SmoothQuant Precision deployed on NVIDIA A100 80GB PCIe.
CogVideoX-5B Video Synthesis (AWQ 4-Bit Activation-Aware) on NVIDIA RTX 4090 24GB GDDR6X VRAM & Throughput Calculator
Exact VRAM memory allocation, dynamic KV-cache requirements, and tensor parallelism slicing for CogVideoX-5B Video Synthesis quantized in AWQ 4-Bit Activation-Aware deployed on NVIDIA RTX 4090 24GB GDDR6X.
CogVideoX-5B Video Synthesis (GPTQ 4-Bit Second-Order) on NVIDIA L40S 48GB Ada Lovelace VRAM & Throughput Calculator
Exact VRAM memory allocation, dynamic KV-cache requirements, and tensor parallelism slicing for CogVideoX-5B Video Synthesis quantized in GPTQ 4-Bit Second-Order deployed on NVIDIA L40S 48GB Ada Lovelace.
CogVideoX-5B Video Synthesis (GGUF Q4_K_M Medium Quant) on AMD Instinct MI300X 192GB VRAM & Throughput Calculator
Exact VRAM memory allocation, dynamic KV-cache requirements, and tensor parallelism slicing for CogVideoX-5B Video Synthesis quantized in GGUF Q4_K_M Medium Quant deployed on AMD Instinct MI300X 192GB.
Whisper Large v3 Audio Speech-to-Text (FP16 Uncompressed Native) on NVIDIA H100 80GB SXM5 VRAM & Throughput Calculator
Exact VRAM memory allocation, dynamic KV-cache requirements, and tensor parallelism slicing for Whisper Large v3 Audio Speech-to-Text quantized in FP16 Uncompressed Native deployed on NVIDIA H100 80GB SXM5.
Whisper Large v3 Audio Speech-to-Text (BF16 Bfloat16 Mixed Precision) on NVIDIA H200 141GB HBM3e VRAM & Throughput Calculator
Exact VRAM memory allocation, dynamic KV-cache requirements, and tensor parallelism slicing for Whisper Large v3 Audio Speech-to-Text quantized in BF16 Bfloat16 Mixed Precision deployed on NVIDIA H200 141GB HBM3e.
Whisper Large v3 Audio Speech-to-Text (FP8 Scaled Native Hopper) on NVIDIA B200 192GB Blackwell VRAM & Throughput Calculator
Exact VRAM memory allocation, dynamic KV-cache requirements, and tensor parallelism slicing for Whisper Large v3 Audio Speech-to-Text quantized in FP8 Scaled Native Hopper deployed on NVIDIA B200 192GB Blackwell.
Whisper Large v3 Audio Speech-to-Text (INT8 SmoothQuant Precision) on NVIDIA A100 80GB PCIe VRAM & Throughput Calculator
Exact VRAM memory allocation, dynamic KV-cache requirements, and tensor parallelism slicing for Whisper Large v3 Audio Speech-to-Text quantized in INT8 SmoothQuant Precision deployed on NVIDIA A100 80GB PCIe.
Whisper Large v3 Audio Speech-to-Text (AWQ 4-Bit Activation-Aware) on NVIDIA RTX 4090 24GB GDDR6X VRAM & Throughput Calculator
Exact VRAM memory allocation, dynamic KV-cache requirements, and tensor parallelism slicing for Whisper Large v3 Audio Speech-to-Text quantized in AWQ 4-Bit Activation-Aware deployed on NVIDIA RTX 4090 24GB GDDR6X.
Whisper Large v3 Audio Speech-to-Text (GPTQ 4-Bit Second-Order) on NVIDIA L40S 48GB Ada Lovelace VRAM & Throughput Calculator
Exact VRAM memory allocation, dynamic KV-cache requirements, and tensor parallelism slicing for Whisper Large v3 Audio Speech-to-Text quantized in GPTQ 4-Bit Second-Order deployed on NVIDIA L40S 48GB Ada Lovelace.
Whisper Large v3 Audio Speech-to-Text (GGUF Q4_K_M Medium Quant) on AMD Instinct MI300X 192GB VRAM & Throughput Calculator
Exact VRAM memory allocation, dynamic KV-cache requirements, and tensor parallelism slicing for Whisper Large v3 Audio Speech-to-Text quantized in GGUF Q4_K_M Medium Quant deployed on AMD Instinct MI300X 192GB.
BGE-M3 Multilingual Embedding 567M (FP16 Uncompressed Native) on NVIDIA H100 80GB SXM5 VRAM & Throughput Calculator
Exact VRAM memory allocation, dynamic KV-cache requirements, and tensor parallelism slicing for BGE-M3 Multilingual Embedding 567M quantized in FP16 Uncompressed Native deployed on NVIDIA H100 80GB SXM5.
BGE-M3 Multilingual Embedding 567M (BF16 Bfloat16 Mixed Precision) on NVIDIA H200 141GB HBM3e VRAM & Throughput Calculator
Exact VRAM memory allocation, dynamic KV-cache requirements, and tensor parallelism slicing for BGE-M3 Multilingual Embedding 567M quantized in BF16 Bfloat16 Mixed Precision deployed on NVIDIA H200 141GB HBM3e.
BGE-M3 Multilingual Embedding 567M (FP8 Scaled Native Hopper) on NVIDIA B200 192GB Blackwell VRAM & Throughput Calculator
Exact VRAM memory allocation, dynamic KV-cache requirements, and tensor parallelism slicing for BGE-M3 Multilingual Embedding 567M quantized in FP8 Scaled Native Hopper deployed on NVIDIA B200 192GB Blackwell.
BGE-M3 Multilingual Embedding 567M (INT8 SmoothQuant Precision) on NVIDIA A100 80GB PCIe VRAM & Throughput Calculator
Exact VRAM memory allocation, dynamic KV-cache requirements, and tensor parallelism slicing for BGE-M3 Multilingual Embedding 567M quantized in INT8 SmoothQuant Precision deployed on NVIDIA A100 80GB PCIe.
BGE-M3 Multilingual Embedding 567M (AWQ 4-Bit Activation-Aware) on NVIDIA RTX 4090 24GB GDDR6X VRAM & Throughput Calculator
Exact VRAM memory allocation, dynamic KV-cache requirements, and tensor parallelism slicing for BGE-M3 Multilingual Embedding 567M quantized in AWQ 4-Bit Activation-Aware deployed on NVIDIA RTX 4090 24GB GDDR6X.
BGE-M3 Multilingual Embedding 567M (GPTQ 4-Bit Second-Order) on NVIDIA L40S 48GB Ada Lovelace VRAM & Throughput Calculator
Exact VRAM memory allocation, dynamic KV-cache requirements, and tensor parallelism slicing for BGE-M3 Multilingual Embedding 567M quantized in GPTQ 4-Bit Second-Order deployed on NVIDIA L40S 48GB Ada Lovelace.
BGE-M3 Multilingual Embedding 567M (GGUF Q4_K_M Medium Quant) on AMD Instinct MI300X 192GB VRAM & Throughput Calculator
Exact VRAM memory allocation, dynamic KV-cache requirements, and tensor parallelism slicing for BGE-M3 Multilingual Embedding 567M quantized in GGUF Q4_K_M Medium Quant deployed on AMD Instinct MI300X 192GB.
LLaVA-NeXT 72B Multimodal Vision (FP16 Uncompressed Native) on NVIDIA H100 80GB SXM5 VRAM & Throughput Calculator
Exact VRAM memory allocation, dynamic KV-cache requirements, and tensor parallelism slicing for LLaVA-NeXT 72B Multimodal Vision quantized in FP16 Uncompressed Native deployed on NVIDIA H100 80GB SXM5.
LLaVA-NeXT 72B Multimodal Vision (BF16 Bfloat16 Mixed Precision) on NVIDIA H200 141GB HBM3e VRAM & Throughput Calculator
Exact VRAM memory allocation, dynamic KV-cache requirements, and tensor parallelism slicing for LLaVA-NeXT 72B Multimodal Vision quantized in BF16 Bfloat16 Mixed Precision deployed on NVIDIA H200 141GB HBM3e.
LLaVA-NeXT 72B Multimodal Vision (FP8 Scaled Native Hopper) on NVIDIA B200 192GB Blackwell VRAM & Throughput Calculator
Exact VRAM memory allocation, dynamic KV-cache requirements, and tensor parallelism slicing for LLaVA-NeXT 72B Multimodal Vision quantized in FP8 Scaled Native Hopper deployed on NVIDIA B200 192GB Blackwell.
LLaVA-NeXT 72B Multimodal Vision (INT8 SmoothQuant Precision) on NVIDIA A100 80GB PCIe VRAM & Throughput Calculator
Exact VRAM memory allocation, dynamic KV-cache requirements, and tensor parallelism slicing for LLaVA-NeXT 72B Multimodal Vision quantized in INT8 SmoothQuant Precision deployed on NVIDIA A100 80GB PCIe.
LLaVA-NeXT 72B Multimodal Vision (AWQ 4-Bit Activation-Aware) on NVIDIA RTX 4090 24GB GDDR6X VRAM & Throughput Calculator
Exact VRAM memory allocation, dynamic KV-cache requirements, and tensor parallelism slicing for LLaVA-NeXT 72B Multimodal Vision quantized in AWQ 4-Bit Activation-Aware deployed on NVIDIA RTX 4090 24GB GDDR6X.
LLaVA-NeXT 72B Multimodal Vision (GPTQ 4-Bit Second-Order) on NVIDIA L40S 48GB Ada Lovelace VRAM & Throughput Calculator
Exact VRAM memory allocation, dynamic KV-cache requirements, and tensor parallelism slicing for LLaVA-NeXT 72B Multimodal Vision quantized in GPTQ 4-Bit Second-Order deployed on NVIDIA L40S 48GB Ada Lovelace.
LLaVA-NeXT 72B Multimodal Vision (GGUF Q4_K_M Medium Quant) on AMD Instinct MI300X 192GB VRAM & Throughput Calculator
Exact VRAM memory allocation, dynamic KV-cache requirements, and tensor parallelism slicing for LLaVA-NeXT 72B Multimodal Vision quantized in GGUF Q4_K_M Medium Quant deployed on AMD Instinct MI300X 192GB.
MiniCPM-V 2.6 8B Omni-Vision (FP16 Uncompressed Native) on NVIDIA H100 80GB SXM5 VRAM & Throughput Calculator
Exact VRAM memory allocation, dynamic KV-cache requirements, and tensor parallelism slicing for MiniCPM-V 2.6 8B Omni-Vision quantized in FP16 Uncompressed Native deployed on NVIDIA H100 80GB SXM5.
MiniCPM-V 2.6 8B Omni-Vision (BF16 Bfloat16 Mixed Precision) on NVIDIA H200 141GB HBM3e VRAM & Throughput Calculator
Exact VRAM memory allocation, dynamic KV-cache requirements, and tensor parallelism slicing for MiniCPM-V 2.6 8B Omni-Vision quantized in BF16 Bfloat16 Mixed Precision deployed on NVIDIA H200 141GB HBM3e.
MiniCPM-V 2.6 8B Omni-Vision (FP8 Scaled Native Hopper) on NVIDIA B200 192GB Blackwell VRAM & Throughput Calculator
Exact VRAM memory allocation, dynamic KV-cache requirements, and tensor parallelism slicing for MiniCPM-V 2.6 8B Omni-Vision quantized in FP8 Scaled Native Hopper deployed on NVIDIA B200 192GB Blackwell.
MiniCPM-V 2.6 8B Omni-Vision (INT8 SmoothQuant Precision) on NVIDIA A100 80GB PCIe VRAM & Throughput Calculator
Exact VRAM memory allocation, dynamic KV-cache requirements, and tensor parallelism slicing for MiniCPM-V 2.6 8B Omni-Vision quantized in INT8 SmoothQuant Precision deployed on NVIDIA A100 80GB PCIe.
MiniCPM-V 2.6 8B Omni-Vision (AWQ 4-Bit Activation-Aware) on NVIDIA RTX 4090 24GB GDDR6X VRAM & Throughput Calculator
Exact VRAM memory allocation, dynamic KV-cache requirements, and tensor parallelism slicing for MiniCPM-V 2.6 8B Omni-Vision quantized in AWQ 4-Bit Activation-Aware deployed on NVIDIA RTX 4090 24GB GDDR6X.
MiniCPM-V 2.6 8B Omni-Vision (GPTQ 4-Bit Second-Order) on NVIDIA L40S 48GB Ada Lovelace VRAM & Throughput Calculator
Exact VRAM memory allocation, dynamic KV-cache requirements, and tensor parallelism slicing for MiniCPM-V 2.6 8B Omni-Vision quantized in GPTQ 4-Bit Second-Order deployed on NVIDIA L40S 48GB Ada Lovelace.
MiniCPM-V 2.6 8B Omni-Vision (GGUF Q4_K_M Medium Quant) on AMD Instinct MI300X 192GB VRAM & Throughput Calculator
Exact VRAM memory allocation, dynamic KV-cache requirements, and tensor parallelism slicing for MiniCPM-V 2.6 8B Omni-Vision quantized in GGUF Q4_K_M Medium Quant deployed on AMD Instinct MI300X 192GB.
InternLM2.5 20B 1M Context (FP16 Uncompressed Native) on NVIDIA H100 80GB SXM5 VRAM & Throughput Calculator
Exact VRAM memory allocation, dynamic KV-cache requirements, and tensor parallelism slicing for InternLM2.5 20B 1M Context quantized in FP16 Uncompressed Native deployed on NVIDIA H100 80GB SXM5.
InternLM2.5 20B 1M Context (BF16 Bfloat16 Mixed Precision) on NVIDIA H200 141GB HBM3e VRAM & Throughput Calculator
Exact VRAM memory allocation, dynamic KV-cache requirements, and tensor parallelism slicing for InternLM2.5 20B 1M Context quantized in BF16 Bfloat16 Mixed Precision deployed on NVIDIA H200 141GB HBM3e.
InternLM2.5 20B 1M Context (FP8 Scaled Native Hopper) on NVIDIA B200 192GB Blackwell VRAM & Throughput Calculator
Exact VRAM memory allocation, dynamic KV-cache requirements, and tensor parallelism slicing for InternLM2.5 20B 1M Context quantized in FP8 Scaled Native Hopper deployed on NVIDIA B200 192GB Blackwell.
InternLM2.5 20B 1M Context (INT8 SmoothQuant Precision) on NVIDIA A100 80GB PCIe VRAM & Throughput Calculator
Exact VRAM memory allocation, dynamic KV-cache requirements, and tensor parallelism slicing for InternLM2.5 20B 1M Context quantized in INT8 SmoothQuant Precision deployed on NVIDIA A100 80GB PCIe.
InternLM2.5 20B 1M Context (AWQ 4-Bit Activation-Aware) on NVIDIA RTX 4090 24GB GDDR6X VRAM & Throughput Calculator
Exact VRAM memory allocation, dynamic KV-cache requirements, and tensor parallelism slicing for InternLM2.5 20B 1M Context quantized in AWQ 4-Bit Activation-Aware deployed on NVIDIA RTX 4090 24GB GDDR6X.
InternLM2.5 20B 1M Context (GPTQ 4-Bit Second-Order) on NVIDIA L40S 48GB Ada Lovelace VRAM & Throughput Calculator
Exact VRAM memory allocation, dynamic KV-cache requirements, and tensor parallelism slicing for InternLM2.5 20B 1M Context quantized in GPTQ 4-Bit Second-Order deployed on NVIDIA L40S 48GB Ada Lovelace.
InternLM2.5 20B 1M Context (GGUF Q4_K_M Medium Quant) on AMD Instinct MI300X 192GB VRAM & Throughput Calculator
Exact VRAM memory allocation, dynamic KV-cache requirements, and tensor parallelism slicing for InternLM2.5 20B 1M Context quantized in GGUF Q4_K_M Medium Quant deployed on AMD Instinct MI300X 192GB.
Baichuan-2 13B Enterprise Chinese (FP16 Uncompressed Native) on NVIDIA H100 80GB SXM5 VRAM & Throughput Calculator
Exact VRAM memory allocation, dynamic KV-cache requirements, and tensor parallelism slicing for Baichuan-2 13B Enterprise Chinese quantized in FP16 Uncompressed Native deployed on NVIDIA H100 80GB SXM5.
Baichuan-2 13B Enterprise Chinese (BF16 Bfloat16 Mixed Precision) on NVIDIA H200 141GB HBM3e VRAM & Throughput Calculator
Exact VRAM memory allocation, dynamic KV-cache requirements, and tensor parallelism slicing for Baichuan-2 13B Enterprise Chinese quantized in BF16 Bfloat16 Mixed Precision deployed on NVIDIA H200 141GB HBM3e.