fix db flow; update readme

This commit is contained in:
bunch 2025-05-08 01:39:08 +08:00
parent bc5a0323dd
commit 32b6cbff81
9 changed files with 125 additions and 16 deletions

View File

@ -1,16 +1,12 @@
# Xplace: An Extremely Fast and Extensible Global Placement Framework
## News 🚀
We are happy to announce that [Xplace 2.0](https://ieeexplore.ieee.org/abstract/document/10373583) is now released. Compared to [Xplace 1.0](https://dl.acm.org/doi/abs/10.1145/3489517.3530485), this version supports the following new features:
We are thrilled to release [Xplace 3.0](https://dl.acm.org/doi/10.1145/3676536.3676803) with timing optimization additional to [Xplace 2.0](https://ieeexplore.ieee.org/abstract/document/10373583) and [Xplace 1.0](https://dl.acm.org/doi/abs/10.1145/3489517.3530485), this version supports the following new features:
- Support deterministic mode with only 5~25% extra GP runtime overhead.
- Implement an extremely fast GPU-accelerated detailed-routability-driven placement algorithm Xplace-Route.
- Integrate with a GPU-accelerated detailed placer and a GPU-accelerated global router [GGR](cpp_to_py/gpugr/README.md).
- Support a superfast **GPU-accelerated place and global route flow**! Input your LEF/DEF, the flow will output the **placement DEF** and the **global routing guide**! [xplace_route_flow.png](img/xplace_route_flow.png)
- Provide benchmark download and preprocess scripts, and three routability evaluation scripts.
- Code refactoring.
- Implement a GPU-accelerated timer.
- Implement an extremely fast GPU-accelerated timing-driven placement algorithm Xplace-Timing.
Please check our [TCAD paper](https://ieeexplore.ieee.org/document/10373583) for more details about **Xplace-Route**!
Please check our [ICCAD paper](https://dl.acm.org/doi/pdf/10.1145/3676536.3676803) for more details about **Xplace-Timing**.
## About Xplace
Xplace is a fast and extensible GPU-accelerated global placement framework developed by the research team supervised by Prof. Evangeline F. Y. Young at The Chinese University of Hong Kong (CUHK). It achieves around 3x speedup per GP iteration compared to DREAMPlace and shows high extensibility.
@ -24,9 +20,11 @@ As shown in the following figure, Xplace framework is built on top of PyTorch an
More details are in the following paper:
Lixin Liu, Bangqi Fu, Martin D. F. Wong, and Evangeline F. Y. Young. "[Xplace: an extremely fast and extensible global placement framework](https://doi.org/10.1145/3489517.3530485)". In Proceedings of the 59th ACM/IEEE Design Automation Conference (DAC '22). Association for Computing Machinery, New York, NY, USA, 1309–1314.
Lixin Liu, Bangqi Fu, Shiju Lin, Jinwei Liu, Evangeline F.Y. Young, Martin D.F. Wong. "[Xplace: An Extremely Fast and Extensible Placement Framework](https://ieeexplore.ieee.org/document/10373583)". In IEEE Transactions on Computer-Aided Design of Integrated Circuits and Systems (TCAD), doi: 10.1109/TCAD.2023.3346291.
Lixin Liu, Bangqi Fu, Martin D. F. Wong, and Evangeline F. Y. Young. "[Xplace: an extremely fast and extensible global placement framework](https://doi.org/10.1145/3489517.3530485)". In Proceedings of the 59th ACM/IEEE Design Automation Conference (DAC '22). Association for Computing Machinery, New York, NY, USA, 1309–1314.
Bangqi Fu, Lixin Liu, Martin D. F. Wong, and Evangeline F. Y. Young. "[Hybrid Modeling and Weighting for Timing-driven Placement with Efficient Calibration]((https://dl.acm.org/doi/10.1145/3676536.3676803))". In Proceedings of the 43rd IEEE/ACM International Conference on Computer-Aided Design (ICCAD '24). Association for Computing Machinery, New York, NY, USA, Article 22, 1–9.
(For the Xplace-NN, please refer to branch [neural](https://github.com/cuhk-eda/Xplace/tree/neural))
@ -95,6 +93,15 @@ python main.py --dataset ispd2015_fix --design_name mgc_fft_1
python main.py --dataset ispd2015_fix --run_all True
```
- To run timing optimization GP + DP flow for ICCAD2015 dataset:
```bash
# only run superblue4
python main.py --dataset iccad2015 --design_name superblue4 --timing_opt True
# run all the designs in iccad2015
python main.py --dataset iccad2015 --run_all True --timing_opt True
```
- To run Routability GP + DP flow for ISPD2015/2018/2019 dataset:
```bash
# run all the designs with routability optimization
@ -128,7 +135,7 @@ Suppose there is a LEF/DEF benchmark named `toy` in `data/raw`, you can use the
python main.py --custom_path lef:data/raw/toy_input.lef,def:data/raw/toy_input.def,design_name:toy,benchmark:test --load_from_raw True --detail_placement True
```
- **Custom JSON Mode**: You can also use the argument `--custom_json` to run multiple LEFs + DEF + Verilog:
- **Custom JSON Mode**: You can also use the argument `--custom_json` to run multiple LEFs + DEF + Verilog + LIBs:
```
python main.py --custom_json examples/examples.json --load_from_raw True --target_density 0.9
```
@ -139,7 +146,13 @@ Please refer to `main.py`.
## 3. Other Features
### 3.1. GPU-accelerated Place and Global Route Flow
### 3.1. Standalone Timer Mode
Setup the design parameters in `tool/timer.py` and run. Put the extracted parasitics file in `spef` option to report the spef timing. An example is given in the `timer.py` file.
```bash
python timer.py
```
### 3.2. GPU-accelerated Place and Global Route Flow
Set `--final_route_eval True` in Python arguments to invoke the internal global router [GGR](cpp_to_py/gpugr/README.md) to run GPU-accelerated PnR flow. The flow will output the **placement DEF** and the **global routing guide** in `./result/exp_id/output`. Besides, GR metrics are reported in the log and recorded in `./result/exp_id/log/route.csv`.
- To run Place and Global Route flow for ISPD2015 dataset:
@ -149,7 +162,7 @@ python main.py --dataset ispd2015_fix --run_all True --load_from_raw True --deta
More details about using GGR in Xplace can be found in [cpp_to_py/gpugr/README.md](cpp_to_py/gpugr/README.md).
### 3.2. Evaluate the Routability of Placement Solution
### 3.3. Evaluate the Routability of Placement Solution
We provide three ways to evaluate the routability of a placement solution:
1. Set `--final_route_eval True` to invoke [GGR](cpp_to_py/gpugr/README.md) to evaluate the placement solution.
@ -160,7 +173,7 @@ We provide three ways to evaluate the routability of a placement solution:
### 3.3. Load Design from Preprocessed File
### 3.4. Load Design from Preprocessed File
The following script will dump the parsed design into a single torch `pt` file so Xplace can load the design from the `pt` file instead of parsing the input file from scratch.
```bash
@ -185,6 +198,14 @@ python main.py --dataset ispd2005 --run_all True --load_from_raw False
## 4. Citation
If you find **Xplace** useful in your research, please consider to cite:
```bibtex
@inproceedings{fu2024xplace_t,
author = {Fu, Bangqi and Liu, Lixin and Wong, Martin D. F. and Young, Evangeline F. Y.},
booktitle = {Proceedings of the 43rd IEEE/ACM International Conference on Computer-Aided Design},
title = {Hybrid Modeling and Weighting for Timing-driven Placement with Efficient Calibration},
year = {2024},
}
@article{xplace_tcad,
author={Liu, Lixin and Fu, Bangqi and Lin, Shiju and Liu, Jinwei and Young, Evangeline F.Y. and Wong, Martin D.F.},
journal={IEEE Transactions on Computer-Aided Design of Integrated Circuits and Systems},
@ -200,7 +221,7 @@ If you find **Xplace** useful in your research, please consider to cite:
}
```
Thanks the authors of [ePlace](https://dl.acm.org/doi/10.1145/2699873), [RePlAce](https://github.com/The-OpenROAD-Project/RePlAce), and [DREAMPlace](https://github.com/limbo018/DREAMPlace) for their great work.
Thanks the authors of [ePlace](https://dl.acm.org/doi/10.1145/2699873), [RePlAce](https://github.com/The-OpenROAD-Project/RePlAce), [DREAMPlace](https://github.com/limbo018/DREAMPlace), [OpenTimer](https://github.com/OpenTimer/OpenTimer), and [GPU-STA](https://ieeexplore.ieee.org/document/9256516) for their great work.
```bibtex
@article{lu2015eplace,
author={Lu, Jingwei and Chen, Pengwen and Chang, Chin-Chih and Sha, Lu and Huang, Dennis Jen-Hsin and Teng, Chin-Chi and Cheng, Chung-Kuan},
@ -222,6 +243,22 @@ Thanks the authors of [ePlace](https://dl.acm.org/doi/10.1145/2699873), [RePlAce
title={DREAMPlace: Deep Learning Toolkit-Enabled GPU Acceleration for Modern VLSI Placement},
year={2021},
}
@article{huang2015opentimer,
author={Huang, Tsung-Wei and Wong, Martin D. F.},
booktitle={2015 IEEE/ACM International Conference on Computer-Aided Design (ICCAD)},
title={OpenTimer: A high-performance timing analysis tool},
year={2015},
}
@article{guo2020gpusta,
author={Guo, Zizheng and Huang, Tsung-Wei and Lin, Yibo},
booktitle={2020 IEEE/ACM International Conference On Computer Aided Design (ICCAD)},
title={GPU-Accelerated Static Timing Analysis},
year={2020},
}
```
## 5. Contact

View File

@ -59,6 +59,7 @@ void Database::load() {
if (setting.BookshelfAux != "") {
setting.Format = "bookshelf";
readBSAux(setting.BookshelfAux, setting.BookshelfPl);
def_read = true;
}
if (setting.LefFile != "") {

View File

@ -14,6 +14,10 @@ std::shared_ptr<gt::GPUTimer> create_gputimer(const py::dict& kwargs,
std::shared_ptr<db::Database> rawdb,
std::shared_ptr<gp::GPDatabase> gpdb,
std::shared_ptr<gt::TimingTorchRawDB> timing_raw_db) {
if (!rawdb->liberty_read) {
throw std::invalid_argument("Liberty file not found. Please check!");
}
std::shared_ptr<gt::GTDatabase> gtdb = std::make_shared<gt::GTDatabase>(rawdb, gpdb, timing_raw_db);
auto sdc = std::make_shared<gt::sdc::SDC>();

View File

@ -7,5 +7,13 @@
"/PDK/asap7/lef/asap7sc7p5t_28_R_1x_220121a.lef",
"/PDK/asap7/lef/asap7sc7p5t_28_SL_1x_220121a.lef"
],
"def": "/designs/top/def/top.def"
"libs": [
"./PDK/asap7/lib/NLDM/asap7sc7p5t_AO_RVT_FF_nldm_211120.lib.gz",
"./PDK/asap7/lib/NLDM/asap7sc7p5t_INVBUF_RVT_FF_nldm_220122.lib.gz",
"./PDK/asap7/lib/NLDM/asap7sc7p5t_OA_RVT_FF_nldm_211120.lib.gz",
"./PDK/asap7/lib/NLDM/asap7sc7p5t_SIMPLE_RVT_FF_nldm_211120.lib.gz",
"./PDK/asap7/lib/NLDM/asap7sc7p5t_SEQ_RVT_FF_nldm_220123.lib"
],
"def": "/designs/top/def/top.def",
"verilog": "/designs/top/def/top.v"
}

View File

@ -50,7 +50,7 @@ def get_option():
parser.add_argument('--visualize_cgmap', type=str2bool, default=False, help='visualize congestion map')
# timing opt params
parser.add_argument('--timing_opt', type=str2bool, default=True, help='perform timing optimization')
parser.add_argument('--timing_opt', type=str2bool, default=False, help='perform timing optimization')
parser.add_argument('--timing_freq', type=int, default=1, help='timing freq')
parser.add_argument('--calibration', type=str2bool, default=True, help='perform timer calibration')
parser.add_argument('--calibration_step', type=float, default=0.1, help='timing calibration step')

View File

@ -117,6 +117,10 @@ class GPUTimer():
self.timer.update_rc(node_lpos, False, True, True)
self.timer.update_timing()
def update_timing_spef(self):
self.timer.update_states()
self.timer.update_rc_spef()
self.timer.update_timing()
def report_timing_slack(self):
time_unit = self.timer.time_unit()

View File

@ -14,6 +14,10 @@ def run_placement_single(args, logger):
def run_placement_all(args, logger):
logger.info("Run all designs in dataset %s." % args.dataset)
place_df = pd.DataFrame(columns=["design", "dp_hpwl", "gp_hpwl", "top5overflow", "overflow", "gp_time", "lg+dp_time", "gp_per_iter", "place_time"])
if args.timing_opt:
place_df = pd.DataFrame(
columns=["design", "dp_hpwl", "gp_hpwl", "top5overflow", "overflow", "gp_time", "lg+dp_time", "gp_per_iter", "place_time", "wns_early_dp", "tns_early_dp", "wns_late_dp", "tns_late_dp"]
)
route_df = pd.DataFrame(columns=["design", "#OvflNets", "GR WL", "GR #Vias", "GR EstShort", "RC Hor", "RC Ver"])
mul_params = sorted(
get_multiple_design_params(args.dataset_root, args.dataset), key=lambda params: params["design_name"]

View File

@ -488,5 +488,7 @@ def run_placement_main_nesterov(args, logger):
logger.info("GP Time: %.4f LG Time: %.4f DP Time: %.4f Total Place Time: %.4f" % (
gp_time, lg_time, dp_time, place_time))
place_metrics = (dp_hpwl, gp_hpwl, top5overflow, overflow, gp_time, dp_time + lg_time, gp_per_iter, place_time)
if args.timing_opt:
place_metrics += (wns_early_dp, tns_early_dp, wns_late_dp, tns_late_dp)
return place_metrics, route_metrics

49
tool/timer.py Normal file
View File

@ -0,0 +1,49 @@
import sys
sys.path.append(os.path.dirname(os.path.dirname(os.path.abspath(__file__))))
from utils import *
from src import Flute, load_dataset, GPUTimer
from main import get_option
def main():
Flute.register(8)
# Read input file
design_name = "example"
params = {
"benchmark": "custom",
"design_name": "test",
"lef": f"{design_name}/NangateOpenCellLibrary.lef",
"lib": f"{design_name}/NangateOpenCellLibrary.lib",
"def": f"{design_name}/example.def",
"verilog": f"{design_name}/example.v",
"sdc": f"{design_name}/example.sdc",
"spef": f"{design_name}/example.spef",
}
args = get_option()
logger = setup_logger(args, sys.argv)
data, rawdb, gpdb = load_dataset(args, logger, params)
device = torch.device(
"cuda:{}".format(args.gpu) if torch.cuda.is_available() else "cpu"
)
data = data.to(device)
data = data.preprocess()
gputimer = GPUTimer(data, rawdb, gpdb, params, args)
# timing analysis for extracted RC network
gputimer.timer.read_spef(params["spef"])
gputimer.update_timing_spef()
wns_early, tns_early, wns_late, tns_late = gputimer.report_timing_slack()
logger.info("SPEf evaluation: wns_early: %.3f, tns_early: %.3f, wns_late: %.3f, tns_late: %.3f" % (wns_early, tns_early, wns_late, tns_late))
# timing analysis for normalized FLUTE RC tree
gputimer.update_timing_eval(data.node_pos)
wns_early, tns_early, wns_late, tns_late = gputimer.report_timing_slack()
logger.info("Flute Tree Evaluation wns_early: %.3f, tns_early: %.3f, wns_late: %.3f, tns_late: %.3f" % (wns_early, tns_early, wns_late, tns_late))
# run main
__name__ == "__main__" and main()