Quantization
Upcoming features
BOS is expanding quantization support across both Vision and LLM/VLM workloads. Planned updates include end-to-end W8A8 LLM compression documentation with llm-compressor, INT8 W8A8 execution for LLM/VLM models, INT8 W8A8 execution for Vision models such as YOLOv8s, ResNet50, DeepLab V3+, PDL, SSR, and FastOFT, and a customer-facing Vision quantization workflow covering calibration, import, execution, validation, profiling, and troubleshooting.
Future documentation will also describe supported recipes, expected artifacts, accuracy validation, known limitations, mixed-precision fallback rules, and how quantized models transition into the BOS-NPU deployment flow.