<?xml version="1.0" encoding="utf-8" standalone="yes"?><rss version="2.0" xmlns:atom="http://www.w3.org/2005/Atom" xmlns:content="http://purl.org/rss/1.0/modules/content/"><channel><title>Optimization on Glory Hinody</title><link>https://glory-hinody.pages.dev/tags/optimization/</link><description>Recent content in Optimization on Glory Hinody</description><generator>Hugo</generator><language>vi</language><lastBuildDate>Tue, 29 Sep 2026 12:00:00 +0700</lastBuildDate><atom:link href="https://glory-hinody.pages.dev/tags/optimization/index.xml" rel="self" type="application/rss+xml"/><item><title>Quantization Int8/FP4: Chạy Model AI Lớn Trên GPU Tài Nguyên Giới Hạn</title><link>https://glory-hinody.pages.dev/posts/quantization-int8-fp4-inference/</link><pubDate>Tue, 29 Sep 2026 12:00:00 +0700</pubDate><guid>https://glory-hinody.pages.dev/posts/quantization-int8-fp4-inference/</guid><description>Phân tích toán học và đo đạc thực nghiệm kỹ thuật nén lượng tử hóa Int8 và FP4 giúp chạy các mô hình Diffusion Transformer 13B–30B trên GPU 16GB không bị tràn bộ nhớ (OOM).</description></item></channel></rss>