# heterogeneity-aware-lowering-and-optimization **Repository Path**: alibaba/heterogeneity-aware-lowering-and-optimization ## Basic Information - **Project Name**: heterogeneity-aware-lowering-and-optimization - **Description**: heterogeneity-aware-lowering-and-optimization - **Primary Language**: Unknown - **License**: Apache-2.0 - **Default Branch**: master - **Homepage**: None - **GVP Project**: No ## Statistics - **Stars**: 0 - **Forks**: 0 - **Created**: 2024-10-31 - **Last Updated**: 2026-10-02 ## Categories & Tags **Categories**: Uncategorized **Tags**: None ## README [](https://opensource.org/licenses/Apache-2.0) [](http://makeapullrequest.com) /badge.svg?branch=master) /badge.svg?branch=master)  HALO =============== **H**eterogeneity-**A**ware **L**owering and **O**ptimization (**HALO**) is a heterogeneous computing acceleration platform based on the compiler technology. It exploits the heterogeneous computing power targeting the deep learning field through an abstract, extendable interface called Open Deep Learning API (**ODLA**). HALO provides a unified Ahead-Of-Time compilation solution, auto tailored for cloud, edge, and IoT scenarios. HALO supports multiple compilation modes. Under the ahead-of-time (AOT) compilation mode, HALO compiles an AI model into the C/C++ code written in the ODLA APIs. The compiled model can be run on any supported platform with the corresponding ODLA runtime liibrary. Plus, HALO is able to compile both host and heterogeneous device code simultaneously. The picture below shows the overall compilation flow:
A broad ODLA ecosystem is supported via the ODLA runtime library set
targeting various heterogeneous accelerators/runtimes:
- [Eigen](http://eigen.tuxfamily.org)
- [Graphcore® IPU](https://www.graphcore.ai)
- [Intel® oneAPI](https://github.com/oneapi-src)
- [Qualcomm® Cloud AI 100](https://www.qualcomm.com/products/cloud-artificial-intelligence)
- [TensorRT™](https://developer.nvidia.com/tensorrt)
- [XNNPACK](https://github.com/google/XNNPACK)
And we welcome new accelerator platforms to join in the ODLA community.
ODLA API Reference can be found
[here](https://alibaba.github.io/heterogeneity-aware-lowering-and-optimization/odla_docs/html/index.html)
and detailed programming guide be coming soon...
## Partners
We appreciate the support of ODLA runtimes from the following partners:
## How to Use HALO
To build HALO, please follow the instructions [here](docs/how_to_build.md) ([查看中文](docs/README_CN.md)).
The workflow of deploying models using HALO includes:
1. Use HALO to compile the model file(s) into an ODLA-based C/C++ source file.
2. Use a C/C++ compiler to compile the generated C/C++ file into an object file.
3. Link the object file, the weight binary, and specific ODLA runtime library together.
### A Simple Example
Let's start with a simple example of MNIST based on
[TensorFlow Tutorial](https://chromium.googlesource.com/external/github.com/tensorflow/tensorflow/+/r0.10/tensorflow/g3doc/tutorials/mnist/beginners/index.md).
The diagram below shows the overall workflow: