{"success":true,"data":[{"id":"fc070c6d-c642-4bf2-b60e-95c3e20d7ca9","name":"RTX PRO 6000 WS 96GB","architecture":"Blackwell","vramGb":96,"memBandwidthGbs":1792,"hourlyPrice":0.83,"cpuPerCoreHr":0.005,"diskPerGbHr":0.0002,"vastAiPrice":1.2,"publicPriceMin":0.84375,"publicPriceMax":1.3125,"liveRentablePrice":1.0738,"offers":[{"gpuCount":1,"hourlyRate":1.0738,"isSlice":false}],"savingsPct":31,"demandStatus":"高用量","demandStatusOverride":null,"totalGpus":6,"availableGpus":1,"availability":"AVAILABLE","rentalMode":"SELF_SERVE","contactEmail":"support.gputw.ai@gmail.com","contactRentalNote":"Email-based long-session rental only.","isActive":true,"pinnedHome":true,"sortOrder":0,"heroImageUrl":"/gpus/rtx-pro-6000.jpg","tagline":"NVIDIA's flagship workstation GPU - 96 GB GDDR7, Blackwell architecture.","taglineZh":"NVIDIA 旗艦工作站 GPU - 96 GB GDDR7、Blackwell 架構。","description":"The NVIDIA RTX PRO 6000 Blackwell Workstation Edition is the most powerful professional desktop GPU ever built. Featuring 96 GB of ECC GDDR7 memory and 24,064 CUDA cores on the full GB202 die, it handles AI inference, 3D rendering, scientific simulation, and video production workflows that no previous workstation GPU could manage in a single card. Its 1.79 TB/s memory bandwidth makes it particularly effective for large-model inference and fine-tuning tasks that exceed the capacity of consumer GPUs.","descriptionZh":"NVIDIA RTX PRO 6000 Blackwell Workstation Edition 是頂級專業桌面 GPU，搭載 96 GB ECC GDDR7 記憶體與 24,064 個 CUDA 核心，適合大型模型推論、LoRA 或完整微調、3D/VFX 渲染、科學模擬與 AI 影片生成。1.79 TB/s 記憶體頻寬與大容量 VRAM 讓它能承載許多消費級 GPU 無法一次放入的模型與資料。若您的工作負載需要單卡高記憶體容量、可靠性與穩定長時間運算，這是 GPUtw 上最完整的工作站級選擇。","chipName":"GB202","processNode":"5 nm (TSMC 4NP)","transistorsBillion":92.2,"releaseDate":"Mar 18, 2025","launchPriceUsd":null,"cudaCores":24064,"tensorCoreGen":5,"tensorCores":752,"rtCoreGen":4,"rtCores":188,"baseClockMhz":1590,"boostClockMhz":2617,"memoryType":"GDDR7 ECC","memoryBusBits":512,"tdpW":600,"fp32Tflops":126,"fp16Tflops":126,"aiTops":4000,"cardCount":1,"perCardMemoryGb":null,"perCardCudaCores":null,"interconnect":null,"idealFor":["Large model inference (70 B+ params)","LLM fine-tuning (LoRA / full)","3D & VFX rendering","Scientific simulation","AI video generation"]},{"id":"0e28d9c6-3156-4ba2-bb99-d5f6679e858e","name":"DGX Spark GB10 4TB","architecture":"GB10 Blackwell","vramGb":128,"memBandwidthGbs":273,"hourlyPrice":0.8125,"cpuPerCoreHr":0.005,"diskPerGbHr":0.0002,"vastAiPrice":null,"publicPriceMin":null,"publicPriceMax":null,"liveRentablePrice":null,"offers":[],"savingsPct":null,"demandStatus":"售罄","demandStatusOverride":null,"totalGpus":1,"availableGpus":0,"availability":"BUSY","rentalMode":"SELF_SERVE","contactEmail":"support.gputw.ai@gmail.com","contactRentalNote":"Email-based long-session rental only.","isActive":true,"pinnedHome":true,"sortOrder":1,"heroImageUrl":"/gpus/dgx-spark.jpg","tagline":"NVIDIA's personal AI supercomputer - GB10 Superchip with 128 GB unified memory.","taglineZh":"NVIDIA 個人 AI 超級電腦 - GB10 Superchip 與 128 GB 統一記憶體。","description":"The NVIDIA DGX Spark is a compact desktop AI supercomputer powered by the GB10 Superchip, combining a 20-core Arm CPU with a Blackwell-generation GPU in a small desktop chassis. Its 128 GB of unified LPDDR5X memory is shared seamlessly between CPU and GPU, enabling models too large for any single-card GPU to run locally. With up to 1,000 AI TOPS and 1 PFLOP at FP4 precision, it is purpose-built for local LLM inference, fine-tuning with quantized weights, and agentic AI workflows that require low-latency GPU access.","descriptionZh":"NVIDIA DGX Spark 是小型桌面 AI 超級電腦，採用 GB10 Superchip，將 20 核 Arm CPU 與 Blackwell 世代 GPU 整合在同一系統中。128 GB LPDDR5X 統一記憶體可由 CPU 與 GPU 共用，適合在本地執行單卡 GPU 放不下的大型模型、長上下文推論、量化微調與 agentic AI 工作流。它的重點不是傳統顯卡式峰值效能，而是以低延遲、低功耗與大記憶體池支援隱私敏感或原型開發場景。","chipName":"GB10 Superchip","processNode":"4 nm (TSMC 4NP)","transistorsBillion":null,"releaseDate":"Mar 2025 (available Oct 2025)","launchPriceUsd":3999,"cudaCores":6144,"tensorCoreGen":5,"tensorCores":null,"rtCoreGen":4,"rtCores":null,"baseClockMhz":0,"boostClockMhz":0,"memoryType":"LPDDR5X (unified CPU+GPU)","memoryBusBits":256,"tdpW":240,"fp32Tflops":5,"fp16Tflops":null,"aiTops":1000,"cardCount":1,"perCardMemoryGb":null,"perCardCudaCores":null,"interconnect":null,"idealFor":["Large model inference (70 B+ at int4)","Local agentic AI workflows","LLM fine-tuning with quantization","Prototyping on large context windows","Privacy-sensitive inference (on-premise)"]},{"id":"c0c092fc-c55f-47be-aa9e-b11d627f73e3","name":"RTX 5090 32GB","architecture":"Blackwell","vramGb":32,"memBandwidthGbs":1792,"hourlyPrice":0.32,"cpuPerCoreHr":0.005,"diskPerGbHr":0.0002,"vastAiPrice":0.4,"publicPriceMin":0.3125,"publicPriceMax":0.84375,"liveRentablePrice":0.5606,"offers":[{"gpuCount":1,"hourlyRate":0.5606,"isSlice":false}],"savingsPct":20,"demandStatus":"高用量","demandStatusOverride":null,"totalGpus":4,"availableGpus":1,"availability":"AVAILABLE","rentalMode":"SELF_SERVE","contactEmail":"support.gputw.ai@gmail.com","contactRentalNote":"Email-based long-session rental only.","isActive":true,"pinnedHome":true,"sortOrder":2,"heroImageUrl":"/gpus/rtx-5090.jpg","tagline":"NVIDIA's fastest consumer GPU - Blackwell architecture with 32 GB GDDR7.","taglineZh":"NVIDIA Blackwell 世代最快消費級 GPU - 32 GB GDDR7。","description":"The GeForce RTX 5090 is the flagship consumer GPU of NVIDIA's Blackwell generation. Built on the same GB202 die as the RTX PRO 6000 WS with 21,760 active CUDA cores and 32 GB of GDDR7 on a 512-bit bus, it delivers the same 1.79 TB/s memory bandwidth at a fraction of the cost. The RTX 5090 offers exceptional performance for AI inference workloads up to 32 GB, diffusion model generation, and gaming, making it the most versatile Blackwell option for general GPU compute.","descriptionZh":"GeForce RTX 5090 是 Blackwell 世代旗艦消費級 GPU，擁有 21,760 個 CUDA 核心、32 GB GDDR7 與 512-bit 記憶體匯流排，頻寬最高達 1.79 TB/s。它非常適合 32 GB VRAM 以內的 AI 推論、影像與影片生成、LoRA 微調、電腦視覺訓練與即時渲染。相較專業卡，RTX 5090 以更低成本提供極高吞吐量，是一般 AI 開發與創作工作流的高效選擇。","chipName":"GB202","processNode":"5 nm (TSMC 4NP)","transistorsBillion":92.2,"releaseDate":"Jan 30, 2025","launchPriceUsd":null,"cudaCores":21760,"tensorCoreGen":5,"tensorCores":680,"rtCoreGen":4,"rtCores":170,"baseClockMhz":2010,"boostClockMhz":2407,"memoryType":"GDDR7","memoryBusBits":512,"tdpW":575,"fp32Tflops":104.8,"fp16Tflops":209.6,"aiTops":3352,"cardCount":1,"perCardMemoryGb":null,"perCardCudaCores":null,"interconnect":null,"idealFor":["AI inference (up to 32 GB models)","Image & video generation","LLM fine-tuning (LoRA)","Computer vision training","Gaming & real-time rendering"]},{"id":"d6170519-d623-47b2-9760-a20ed4b64db6","name":"RTX 3090 24GB","architecture":"Ampere","vramGb":24,"memBandwidthGbs":936,"hourlyPrice":0.12,"cpuPerCoreHr":0.005,"diskPerGbHr":0.0002,"vastAiPrice":0.15,"publicPriceMin":0.125,"publicPriceMax":0.375,"liveRentablePrice":0.2247,"offers":[{"gpuCount":1,"hourlyRate":0.2247,"isSlice":false}],"savingsPct":20,"demandStatus":"高用量","demandStatusOverride":null,"totalGpus":5,"availableGpus":1,"availability":"AVAILABLE","rentalMode":"SELF_SERVE","contactEmail":"support.gputw.ai@gmail.com","contactRentalNote":"Email-based long-session rental only.","isActive":true,"pinnedHome":true,"sortOrder":3,"heroImageUrl":"/gpus/rtx-3090.jpg","tagline":"Flagship Ampere consumer GPU - 24 GB GDDR6X, 10,496 CUDA cores.","taglineZh":"Ampere 旗艦消費級 GPU - 24 GB GDDR6X、10,496 CUDA 核心。","description":"The GeForce RTX 3090 was NVIDIA's highest-end consumer GPU of the Ampere generation. With 24 GB of GDDR6X on a 384-bit bus delivering 936 GB/s of bandwidth, it remains a capable workhorse for AI training and inference tasks that fit in 24 GB. Its 10,496 CUDA cores and third-generation Tensor Cores provide strong FP16 and TF32 throughput, making it a cost-effective option for fine-tuning models up to about 13 B parameters with quantization, stable diffusion image generation, and traditional deep learning workloads.","descriptionZh":"GeForce RTX 3090 具備 24 GB GDDR6X、936 GB/s 記憶體頻寬與第三代 Tensor Cores，至今仍是 AI 訓練與推論的實用工作馬。它適合 24 GB VRAM 內的模型微調、小批次 LLM 推論、Stable Diffusion / ComfyUI、電腦視覺訓練與影片處理。對需要比入門 GPU 更大記憶體、但不需要最新 Blackwell 成本的使用者，RTX 3090 是很有性價比的選擇。","chipName":"GA102","processNode":"8 nm (Samsung)","transistorsBillion":28.3,"releaseDate":"Sep 24, 2020","launchPriceUsd":1499,"cudaCores":10496,"tensorCoreGen":3,"tensorCores":328,"rtCoreGen":2,"rtCores":82,"baseClockMhz":1395,"boostClockMhz":1695,"memoryType":"GDDR6X","memoryBusBits":384,"tdpW":350,"fp32Tflops":35.6,"fp16Tflops":71,"aiTops":null,"cardCount":1,"perCardMemoryGb":null,"perCardCudaCores":null,"interconnect":null,"idealFor":["Fine-tuning up to 13 B (int4)","Stable Diffusion & ComfyUI","Small-batch LLM inference","Computer vision training","Video processing"]},{"id":"750748eb-8229-48f5-8d25-a56899cddd52","name":"RTX 3090 x2 48GB NVLink","architecture":"Ampere","vramGb":48,"memBandwidthGbs":1872,"hourlyPrice":0.12,"cpuPerCoreHr":0.005,"diskPerGbHr":0.0002,"vastAiPrice":0.3,"publicPriceMin":0.25,"publicPriceMax":0.75,"liveRentablePrice":null,"offers":[],"savingsPct":60,"demandStatus":"售罄","demandStatusOverride":null,"totalGpus":2,"availableGpus":0,"availability":"BUSY","rentalMode":"SELF_SERVE","contactEmail":"support.gputw.ai@gmail.com","contactRentalNote":"Email-based long-session rental only.","isActive":true,"pinnedHome":true,"sortOrder":4,"heroImageUrl":"/gpus/rtx-3090.jpg","tagline":"Two RTX 3090 GPUs bridged by NVLink - 48 GB of combined Ampere VRAM.","taglineZh":"兩張 RTX 3090 透過 NVLink 連接 - 48 GB Ampere VRAM。","description":"This configuration pairs two GeForce RTX 3090 cards via NVIDIA NVLink, presenting 48 GB of aggregate GDDR6X VRAM as a unified address space. NVLink provides up to 112.5 GB/s of bidirectional bandwidth between the cards, dramatically outpacing PCIe for inter-GPU communication. The combined 20,992 CUDA cores and 1.87 TB/s of total memory bandwidth make this config well-suited for model-parallel training of LLMs that fit in 48 GB, as well as batch inference workloads that benefit from high aggregate throughput.","descriptionZh":"此配置使用兩張 GeForce RTX 3090 並透過 NVIDIA NVLink 連接，提供 48 GB 總 GDDR6X VRAM 與更快的卡間通訊。NVLink 對模型平行、批次推論與需要跨 GPU 傳輸的工作負載特別有利，比只靠 PCIe 更適合多卡協作。對於可切分到 48 GB 記憶體內的 LLM 訓練、3D 渲染、資料科學管線與高吞吐推論，這是成本與容量平衡良好的 Ampere 多卡方案。","chipName":"GA102 x 2","processNode":"8 nm (Samsung)","transistorsBillion":56.6,"releaseDate":"Sep 24, 2020","launchPriceUsd":null,"cudaCores":20992,"tensorCoreGen":3,"tensorCores":656,"rtCoreGen":2,"rtCores":164,"baseClockMhz":1395,"boostClockMhz":1695,"memoryType":"GDDR6X","memoryBusBits":768,"tdpW":700,"fp32Tflops":71.2,"fp16Tflops":142.4,"aiTops":null,"cardCount":2,"perCardMemoryGb":24,"perCardCudaCores":10496,"interconnect":"NVLink 3.0 (112.5 GB/s bidirectional)","idealFor":["Model-parallel LLM training (<= 48 GB)","Batch inference with large context","3D rendering farms","Data science pipelines"]},{"id":"4420e1e5-2c1f-4ac3-a0c3-05c84820d779","name":"H100","architecture":"Hopper","vramGb":80,"memBandwidthGbs":3350,"hourlyPrice":1.67,"cpuPerCoreHr":0.005,"diskPerGbHr":0.0002,"vastAiPrice":2.39,"publicPriceMin":1.65625,"publicPriceMax":2.34375,"liveRentablePrice":null,"offers":[],"savingsPct":30,"demandStatus":"售罄","demandStatusOverride":null,"totalGpus":0,"availableGpus":0,"availability":"NO_HARDWARE","rentalMode":"SELF_SERVE","contactEmail":"support.gputw.ai@gmail.com","contactRentalNote":"Email-based long-session rental only.","isActive":true,"pinnedHome":true,"sortOrder":5,"heroImageUrl":"/gpus/h200.jpg","tagline":"Hopper data center GPU - 80 GB HBM3 for LLM inference, training, and HPC.","taglineZh":"Hopper 資料中心 GPU - 80 GB HBM3，適合 LLM 推論、訓練與 HPC。","description":"The NVIDIA H100 Tensor Core GPU is a Hopper-generation data center accelerator built for large language models, generative AI, recommender systems, and high-performance computing. The SXM 80 GB variant provides 80 GB of HBM3 memory, 3.35 TB/s of memory bandwidth, fourth-generation Tensor Cores, FP8 acceleration, and MIG partitioning for multi-tenant workloads. It remains one of the most widely supported accelerators for production LLM inference, fine-tuning, distributed training, and frameworks that are already tuned for Hopper.","descriptionZh":"NVIDIA H100 Tensor Core GPU 是 Hopper 世代資料中心加速器，面向大型語言模型、生成式 AI、推薦系統與高效能運算。SXM 80 GB 版本提供 80 GB HBM3 記憶體、3.35 TB/s 記憶體頻寬、第四代 Tensor Cores、FP8 加速與 MIG 分割能力。它仍是生產級 LLM 推論、模型微調、分散式訓練，以及已針對 Hopper 最佳化框架中最常見且支援成熟的 GPU 之一。","chipName":"GH100","processNode":"4 nm (TSMC 4N)","transistorsBillion":null,"releaseDate":"Mar 2022","launchPriceUsd":30000,"cudaCores":16896,"tensorCoreGen":4,"tensorCores":528,"rtCoreGen":null,"rtCores":null,"baseClockMhz":0,"boostClockMhz":0,"memoryType":"HBM3","memoryBusBits":5120,"tdpW":700,"fp32Tflops":67,"fp16Tflops":1979,"aiTops":3958,"cardCount":1,"perCardMemoryGb":null,"perCardCudaCores":null,"interconnect":"NVLink 4 (platform dependent)","idealFor":["Production LLM inference","LLM fine-tuning","Distributed deep learning","RAG and embedding workloads","HPC and simulation"]},{"id":"b0fe25b9-7033-407a-9f72-a2ec27493385","name":"H200","architecture":"Hopper","vramGb":141,"memBandwidthGbs":4800,"hourlyPrice":2.8125,"cpuPerCoreHr":0.005,"diskPerGbHr":0.0002,"vastAiPrice":3.28125,"publicPriceMin":2.8125,"publicPriceMax":3.4375,"liveRentablePrice":null,"offers":[],"savingsPct":14,"demandStatus":"售罄","demandStatusOverride":null,"totalGpus":0,"availableGpus":0,"availability":"NO_HARDWARE","rentalMode":"SELF_SERVE","contactEmail":"support.gputw.ai@gmail.com","contactRentalNote":"Email-based long-session rental only.","isActive":true,"pinnedHome":false,"sortOrder":6,"heroImageUrl":"/gpus/h200.jpg","tagline":"NVIDIA Hopper data center GPU - 141 GB HBM3e and 4.8 TB/s memory bandwidth.","taglineZh":"NVIDIA Hopper 資料中心 GPU - 141 GB HBM3e、4.8 TB/s 頻寬。","description":"The NVIDIA H200 Tensor Core GPU is a Hopper-generation data center accelerator built for generative AI and high-performance computing workloads that need very large, very fast memory. With 141 GB of HBM3e memory and 4.8 TB/s of memory bandwidth, H200 is especially strong for large language model inference, retrieval-augmented generation, high-throughput batch inference, and memory-bound HPC applications.","descriptionZh":"NVIDIA H200 Tensor Core GPU 是 Hopper 世代資料中心加速器，面向需要超大且高速記憶體的生成式 AI 與高效能運算工作負載。141 GB HBM3e 與 4.8 TB/s 記憶體頻寬讓它特別適合大型語言模型推論、RAG、高吞吐批次推論、記憶體頻寬受限的 HPC 應用與大型模型微調。若工作負載主要受 VRAM 容量或頻寬限制，H200 是高階選項。","chipName":"GH100","processNode":"4 nm (TSMC 4N)","transistorsBillion":null,"releaseDate":"Nov 2023","launchPriceUsd":null,"cudaCores":16896,"tensorCoreGen":4,"tensorCores":528,"rtCoreGen":null,"rtCores":null,"baseClockMhz":0,"boostClockMhz":0,"memoryType":"HBM3e","memoryBusBits":5120,"tdpW":700,"fp32Tflops":67,"fp16Tflops":1979,"aiTops":3958,"cardCount":1,"perCardMemoryGb":null,"perCardCudaCores":null,"interconnect":"NVLink / NVSwitch platform dependent","idealFor":["Large language model inference","Retrieval-augmented generation","High-throughput batch inference","Memory-bound HPC workloads","Large model fine-tuning"]},{"id":"a9e45eda-9495-4229-8026-dab102648671","name":"B200","architecture":"Blackwell","vramGb":180,"memBandwidthGbs":8000,"hourlyPrice":3.75,"cpuPerCoreHr":0.005,"diskPerGbHr":0.0002,"vastAiPrice":4.3,"publicPriceMin":3.75,"publicPriceMax":4.6875,"liveRentablePrice":null,"offers":[],"savingsPct":13,"demandStatus":"售罄","demandStatusOverride":null,"totalGpus":0,"availableGpus":0,"availability":"NO_HARDWARE","rentalMode":"SELF_SERVE","contactEmail":"support.gputw.ai@gmail.com","contactRentalNote":"Email-based long-session rental only.","isActive":true,"pinnedHome":false,"sortOrder":7,"heroImageUrl":"/api/gpu-images/1169b5867c1dbd4b1e2def12.jpg","tagline":"Blackwell data center GPU - 180 GB HBM3e for large-scale AI inference and training.","taglineZh":"Blackwell 資料中心 GPU - 180 GB HBM3e，適合大規模 AI 推論與訓練。","description":"The NVIDIA B200 is a Blackwell data center GPU designed for large-scale AI training and inference systems. Each B200 SXM GPU provides 180 GB of HBM3e memory and up to 8 TB/s of memory bandwidth, making it highly effective for frontier LLM inference, long-context serving, high-throughput batch inference, large model fine-tuning, and memory-bound HPC. B200 GPUs are commonly deployed in HGX or DGX systems with NVLink / NVSwitch fabrics, where an 8-GPU node provides 1.44 TB of total GPU memory.","descriptionZh":"NVIDIA B200 是 Blackwell 世代資料中心 GPU，設計用於大規模 AI 訓練與推論系統。單張 B200 SXM GPU 提供 180 GB HBM3e 記憶體與最高 8 TB/s 記憶體頻寬，適合大型語言模型推論、長上下文服務、高吞吐批次推論、大型模型微調與記憶體頻寬受限的 HPC 工作負載。B200 通常部署在 HGX 或 DGX 多 GPU 系統中，透過 NVLink / NVSwitch 串接，8-GPU 節點可提供 1.44 TB 總 GPU 記憶體。","chipName":"GB200","processNode":"4 nm (TSMC 4NP)","transistorsBillion":208,"releaseDate":"2024","launchPriceUsd":null,"cudaCores":null,"tensorCoreGen":5,"tensorCores":null,"rtCoreGen":null,"rtCores":null,"baseClockMhz":0,"boostClockMhz":0,"memoryType":"HBM3e","memoryBusBits":8192,"tdpW":1000,"fp32Tflops":75,"fp16Tflops":2250,"aiTops":null,"cardCount":1,"perCardMemoryGb":null,"perCardCudaCores":null,"interconnect":"NVLink 5 / NVSwitch platform dependent","idealFor":["Large language model inference","Long-context serving","High-throughput batch inference","Large model fine-tuning","Memory-bound HPC workloads"]},{"id":"58640237-4f50-49aa-9288-9319a54d8a68","name":"V100 x8 NVLink 128GB","architecture":"Volta","vramGb":128,"memBandwidthGbs":7200,"hourlyPrice":1.08,"cpuPerCoreHr":0.005,"diskPerGbHr":0.0002,"vastAiPrice":1.36,"publicPriceMin":1,"publicPriceMax":1.625,"liveRentablePrice":null,"offers":[],"savingsPct":21,"demandStatus":"售罄","demandStatusOverride":null,"totalGpus":0,"availableGpus":0,"availability":"NO_HARDWARE","rentalMode":"SELF_SERVE","contactEmail":"support.gputw.ai@gmail.com","contactRentalNote":"Email-based long-session rental only.","isActive":true,"pinnedHome":true,"sortOrder":8,"heroImageUrl":"/api/gpu-images/6618bb3565c31e62801574b9.webp","tagline":"Eight V100 GPUs in a full DGX-1 NVLink mesh - 128 GB HBM2, 7.2 TB/s aggregate.","taglineZh":"八張 V100 的 DGX-1 類 NVLink 拓撲 - 128 GB HBM2。","description":"This eight-card V100 NVLink configuration mirrors the interconnect topology of the NVIDIA DGX-1, connecting all eight cards in a hybrid cube-mesh NVLink topology with up to 300 GB/s of bidirectional bandwidth per GPU. With 128 GB of total HBM2 memory and 40,960 CUDA cores across all cards, this remains a proven platform for multi-GPU deep learning training and large-scale HPC workloads. The V100's first-generation Tensor Cores were the original hardware accelerators for mixed-precision training, and mature framework support across PyTorch, TensorFlow, and JAX makes this configuration highly compatible.","descriptionZh":"V100 x8 NVLink 配置接近 NVIDIA DGX-1 的混合 cube-mesh 拓撲，提供 128 GB 總 HBM2 記憶體、40,960 個 CUDA 核心與成熟的多 GPU 深度學習相容性。雖然架構較舊，但 V100 是 Tensor Core 訓練生態的基礎平台，PyTorch、TensorFlow、JAX、DDP 與許多 HPC 程式碼都支援良好。它適合資料平行訓練、大規模科學運算、高吞吐批次推論與需要穩定多卡環境的工作。","chipName":"GV100 x 8","processNode":"12 nm (TSMC)","transistorsBillion":168.8,"releaseDate":"Jun 2017","launchPriceUsd":null,"cudaCores":40960,"tensorCoreGen":1,"tensorCores":5120,"rtCoreGen":null,"rtCores":null,"baseClockMhz":1245,"boostClockMhz":1530,"memoryType":"HBM2","memoryBusBits":32768,"tdpW":2400,"fp32Tflops":125.6,"fp16Tflops":1004.8,"aiTops":null,"cardCount":8,"perCardMemoryGb":16,"perCardCudaCores":5120,"interconnect":"NVLink 2.0 mesh (300 GB/s per GPU bidirectional)","idealFor":["Multi-GPU LLM training","Large-scale scientific HPC","Data-parallel deep learning","High-throughput batch inference"]},{"id":"71eba9e3-be6c-4631-a9ea-3aa079ceb5d4","name":"RTX 4090 24GB","architecture":"Ada Lovelace","vramGb":24,"memBandwidthGbs":1008,"hourlyPrice":0.28,"cpuPerCoreHr":0.005,"diskPerGbHr":0.0002,"vastAiPrice":0.35,"publicPriceMin":0.21875,"publicPriceMax":0.375,"liveRentablePrice":null,"offers":[],"savingsPct":20,"demandStatus":"售罄","demandStatusOverride":null,"totalGpus":0,"availableGpus":0,"availability":"NO_HARDWARE","rentalMode":"SELF_SERVE","contactEmail":"support.gputw.ai@gmail.com","contactRentalNote":"Email-based long-session rental only.","isActive":true,"pinnedHome":false,"sortOrder":9,"heroImageUrl":"/gpus/rtx-4090.jpg","tagline":"Ada Lovelace flagship consumer GPU - 24 GB GDDR6X, 16,384 CUDA cores.","taglineZh":"Ada Lovelace 旗艦消費級 GPU - 24 GB GDDR6X、16,384 CUDA 核心。","description":"The GeForce RTX 4090 is NVIDIA's flagship Ada Lovelace consumer GPU. With 24 GB of GDDR6X memory, 16,384 CUDA cores, fourth-generation Tensor Cores, third-generation RT Cores, and 1,008 GB/s of memory bandwidth, it remains a strong option for AI inference, image generation, video workflows, rendering, and deep learning workloads that fit in 24 GB of VRAM.","descriptionZh":"GeForce RTX 4090 是 NVIDIA Ada Lovelace 世代旗艦消費級 GPU，搭載 24 GB GDDR6X、16,384 個 CUDA 核心、第四代 Tensor Cores、第三代 RT Cores 與 1,008 GB/s 記憶體頻寬。它適合 24 GB VRAM 以內的 AI 推論、Stable Diffusion 和影像生成、量化微調、3D 渲染、影片工作流與電腦視覺訓練。對需要高效能但不需要 32 GB 以上 VRAM 的使用者，RTX 4090 仍是非常強的選擇。","chipName":"AD102","processNode":"4 nm (TSMC 4N)","transistorsBillion":76.3,"releaseDate":"Oct 12, 2022","launchPriceUsd":1599,"cudaCores":16384,"tensorCoreGen":4,"tensorCores":512,"rtCoreGen":3,"rtCores":128,"baseClockMhz":2235,"boostClockMhz":2520,"memoryType":"GDDR6X","memoryBusBits":384,"tdpW":450,"fp32Tflops":82.6,"fp16Tflops":165.2,"aiTops":1321,"cardCount":1,"perCardMemoryGb":null,"perCardCudaCores":null,"interconnect":null,"idealFor":["AI inference up to 24 GB models","Stable Diffusion and image generation","LLM fine-tuning with quantization","3D rendering and video workflows","Computer vision training"]},{"id":"3c96b7db-77aa-4278-90aa-dc27ecb857a1","name":"RTX A5000","architecture":"Ampere","vramGb":24,"memBandwidthGbs":768,"hourlyPrice":0.25,"cpuPerCoreHr":0.005,"diskPerGbHr":0.0002,"vastAiPrice":0.203125,"publicPriceMin":0.125,"publicPriceMax":0.21875,"liveRentablePrice":null,"offers":[],"savingsPct":-23,"demandStatus":"售罄","demandStatusOverride":null,"totalGpus":0,"availableGpus":0,"availability":"NO_HARDWARE","rentalMode":"SELF_SERVE","contactEmail":"support.gputw.ai@gmail.com","contactRentalNote":"Email-based long-session rental only.","isActive":true,"pinnedHome":false,"sortOrder":12,"heroImageUrl":"/gpus/rtx-a5000.jpg","tagline":"Professional Ampere GPU - 24 GB GDDR6 ECC, 8,192 CUDA cores.","taglineZh":"專業 Ampere 工作站 GPU - 24 GB GDDR6 ECC、8,192 CUDA 核心。","description":"The NVIDIA RTX A5000 is a professional Ampere-generation workstation GPU with 24 GB of GDDR6 ECC memory, 8,192 CUDA cores, third-generation Tensor Cores, and second-generation RT Cores. It is a balanced option for rendering, AI development, simulation, computer vision, and data science workloads that benefit from workstation-class memory reliability and NVLink support.","descriptionZh":"NVIDIA RTX A5000 是 Ampere 世代專業工作站 GPU，具備 24 GB GDDR6 ECC 記憶體、8,192 個 CUDA 核心、第三代 Tensor Cores 與第二代 RT Cores。它適合專業渲染、AI 開發、模擬、電腦視覺、資料科學 notebook 與需要工作站級記憶體可靠性的任務。相較消費級 GPU，A5000 更重視穩定性、ECC 記憶體與專業軟體工作流。","chipName":"GA102","processNode":"8 nm (Samsung)","transistorsBillion":null,"releaseDate":"Apr 2021","launchPriceUsd":null,"cudaCores":8192,"tensorCoreGen":3,"tensorCores":256,"rtCoreGen":2,"rtCores":64,"baseClockMhz":1170,"boostClockMhz":1695,"memoryType":"GDDR6 ECC","memoryBusBits":384,"tdpW":230,"fp32Tflops":27.8,"fp16Tflops":null,"aiTops":222,"cardCount":1,"perCardMemoryGb":null,"perCardCudaCores":null,"interconnect":"NVLink 112.5 GB/s bidirectional","idealFor":["Professional rendering","Computer vision training","AI development","Simulation workloads","Data science notebooks"]},{"id":"30581c0d-8bd9-49f3-b0bc-4f64d46a0615","name":"V100 16GB","architecture":"Volta","vramGb":16,"memBandwidthGbs":900,"hourlyPrice":0.12,"cpuPerCoreHr":0.005,"diskPerGbHr":0.0002,"vastAiPrice":0.15,"publicPriceMin":0.125,"publicPriceMax":0.1875,"liveRentablePrice":null,"offers":[],"savingsPct":20,"demandStatus":"售罄","demandStatusOverride":null,"totalGpus":0,"availableGpus":0,"availability":"NO_HARDWARE","rentalMode":"SELF_SERVE","contactEmail":"support.gputw.ai@gmail.com","contactRentalNote":"Email-based long-session rental only.","isActive":true,"pinnedHome":false,"sortOrder":13,"heroImageUrl":"/api/gpu-images/9001f96e175e208aba611162.webp","tagline":"The GPU that launched the Tensor Core era - 16 GB HBM2, first-gen Tensor Cores.","taglineZh":"開啟 Tensor Core 時代的經典 GPU - 16 GB HBM2。","description":"The NVIDIA Tesla V100 SXM2 introduced first-generation Tensor Cores, enabling mixed-precision training that became the standard technique for large-scale deep learning. With 5,120 CUDA cores, 640 Tensor Cores, and 16 GB of HBM2 delivering 900 GB/s of bandwidth, it remains a reliable platform for training and inference tasks up to 16 GB, particularly workloads with mature codebases tuned for V100 compatibility. Its high-bandwidth HBM2 memory makes it efficient for bandwidth-bound operations despite its age.","descriptionZh":"NVIDIA Tesla V100 SXM2 是第一代 Tensor Core GPU，支援讓深度學習普及的混合精度訓練。它具備 5,120 個 CUDA 核心、640 個 Tensor Cores、16 GB HBM2 與 900 GB/s 記憶體頻寬，適合 BERT、ResNet、傳統 Transformer、科學計算、資料前處理與 16 GB 以內的推論任務。V100 的優勢在於成熟框架相容性與高頻寬記憶體，對舊版研究程式碼尤其友善。","chipName":"GV100","processNode":"12 nm (TSMC)","transistorsBillion":21.1,"releaseDate":"Jun 2017","launchPriceUsd":10000,"cudaCores":5120,"tensorCoreGen":1,"tensorCores":640,"rtCoreGen":null,"rtCores":null,"baseClockMhz":1245,"boostClockMhz":1530,"memoryType":"HBM2","memoryBusBits":4096,"tdpW":300,"fp32Tflops":15.7,"fp16Tflops":125.4,"aiTops":null,"cardCount":1,"perCardMemoryGb":null,"perCardCudaCores":null,"interconnect":null,"idealFor":["Classic deep learning (BERT, ResNet, Transformer)","LLM inference (<= 16 GB models)","Scientific computing","Data preprocessing pipelines"]}],"error":null}