ai-agent-book 精选快照(<2MB 代码与文档,来自 github.com/bojieli/ai-agent-book)
Build latest book artifacts / build (push) Canceled after 0s
dependency resolution / resolve (3.11) (push) Canceled after 0s
dependency resolution / resolve (3.13) (push) Canceled after 0s
deploy-pages / build (push) Canceled after 0s
deploy-pages / deploy (push) Canceled after 0s
i18n consistency check / check (push) Canceled after 0s
provider adoption tests / test (chapter2/context-compression) (push) Canceled after 0s
provider adoption tests / test (chapter2/prompt-injection) (push) Canceled after 0s
provider adoption tests / test (chapter2/system-hint) (push) Canceled after 0s
provider adoption tests / test (chapter3/log-sanitization) (push) Canceled after 0s
web-search-agent tests / test (push) Canceled after 0s
web-search-agent tests / agentbook (push) Canceled after 0s
Build latest book artifacts / build (push) Canceled after 0s
dependency resolution / resolve (3.11) (push) Canceled after 0s
dependency resolution / resolve (3.13) (push) Canceled after 0s
deploy-pages / build (push) Canceled after 0s
deploy-pages / deploy (push) Canceled after 0s
i18n consistency check / check (push) Canceled after 0s
provider adoption tests / test (chapter2/context-compression) (push) Canceled after 0s
provider adoption tests / test (chapter2/prompt-injection) (push) Canceled after 0s
provider adoption tests / test (chapter2/system-hint) (push) Canceled after 0s
provider adoption tests / test (chapter3/log-sanitization) (push) Canceled after 0s
web-search-agent tests / test (push) Canceled after 0s
web-search-agent tests / agentbook (push) Canceled after 0s
This commit is contained in:
@@ -0,0 +1,167 @@
|
||||
# Cursor Chat: ai-agent-book
|
||||
|
||||
## Metadata
|
||||
- **Project**: ai-agent-book
|
||||
- **Path**: `/Users/boj`
|
||||
- **Date**: 2025-10-02 19:17:25
|
||||
- **Session ID**: `e83eea6b-7378-4571-9e83-241f98f6befc`
|
||||
|
||||
## Conversation
|
||||
|
||||
### 👤 You
|
||||
|
||||
#!/bin/bash
|
||||
|
||||
#SBATCH --job-name=TrainGP
|
||||
#SBATCH --mail-user=user@gmail.com
|
||||
#SBATCH --mail-type=ALL
|
||||
#SBATCH --nodes=1 # Total number of nodes requested
|
||||
#SBATCH --ntasks-per-node=8 # Total number of task requested
|
||||
#SBATCH --cpus-per-task=8 # Total number of cores requested
|
||||
#SBATCH --mem=1000G
|
||||
#SBATCH --time=96:00:00 # Time limit (hh:mm:ss)
|
||||
#SBATCH --gpus-per-node=8
|
||||
#SBATCH --output=./slurm_logs/train_%A_%a.out
|
||||
#SBATCH --output=./slurm_logs/train_%A_%a.err
|
||||
#SBATCH --partition=gpu
|
||||
|
||||
LR=1e-7
|
||||
save_every=1
|
||||
save_model=False # disable running saving. one checkpoint ~30GB
|
||||
|
||||
CKPT_NAME="YOUR_MODEL_PATH" # official init model: tianzhechu/GP-L-Init
|
||||
PORT=$((RANDOM % 10000 + 1000))
|
||||
|
||||
# download from our huggingface dataset repo tianzhechu/SFTvsRL_Data
|
||||
ROUTE_INFO="YOUR_ROUTE_INFO_PATH" # .json
|
||||
GPS_TO_PANO="YOUR_GPS_TO_PANO_MAPPING_PATH" # .pkl
|
||||
STREETVIEWS="YOUR_STREETVIEWS_PATH" # folder of images
|
||||
|
||||
DS_SKIP_CUDA_CHECK=1 TOKENIZERS_PARALLELISM=false \
|
||||
accelerate launch \
|
||||
--config_file scripts/config_zero2_8gpu.yaml \
|
||||
--main_process_port ${PORT} -m rl.launcher \
|
||||
-f rl/configs/llama_virl_vl.yaml \
|
||||
--output_dir=train_ckpt/virl_vl/ \
|
||||
--optimizer_config.init_lr=${LR} \
|
||||
--optimizer_config.lr_max_steps=20 \
|
||||
--prompt_config.enable_verification=True \
|
||||
--num_updates=15 \
|
||||
--run_name=virl_vl_training \
|
||||
--num_steps=256 \
|
||||
--model_path=${CKPT_NAME} \
|
||||
--save_ckpt=${save_model} \
|
||||
--save_every=${save_every} \
|
||||
--env_config.route_info_path=${ROUTE_INFO} \
|
||||
--env_config.platform_cfg.OFFLINE.PANORAMA_DIR=${STREETVIEWS} \
|
||||
--env_config.platform_cfg.OFFLINE.GPS_TO_PANO_PATH=${GPS_TO_PANO}
|
||||
|
||||
--
|
||||
How many GPUs do I need to run this experiment
|
||||
|
||||
---
|
||||
|
||||
### 🤖 Cursor Assistant
|
||||
|
||||
Based on the SLURM script, you need **8 GPUs** to run this experiment.
|
||||
|
||||
This is specified by the line:
|
||||
```bash
|
||||
#SBATCH --gpus-per-node=8
|
||||
```
|
||||
|
||||
Additional resource requirements:
|
||||
- **1 node** (single machine)
|
||||
- **8 GPUs** on that node
|
||||
- **1TB of RAM** (`--mem=1000G`)
|
||||
- **64 CPU cores** total (8 tasks × 8 CPUs per task)
|
||||
- **96 hours** of runtime
|
||||
|
||||
The script also confirms this with the accelerate configuration file `config_zero2_8gpu.yaml`, which indicates it's using DeepSpeed ZeRO-2 optimization distributed across 8 GPUs.
|
||||
|
||||
**Important notes:**
|
||||
- The model checkpoints are ~30GB each (according to the comment)
|
||||
- This appears to be a vision-language RL training job for a navigation model
|
||||
- You'll need access to a GPU partition with nodes that have 8 GPUs available (typically A100 or H100 nodes)
|
||||
- Make sure you replace the placeholder paths (`YOUR_MODEL_PATH`, `YOUR_ROUTE_INFO_PATH`, etc.) with actual paths before running
|
||||
|
||||
---
|
||||
|
||||
### 👤 You
|
||||
|
||||
File "/root/miniconda3/lib/python3.13/site-packages/cv2/__init__.py", line 181, in <module>
|
||||
bootstrap()
|
||||
~~~~~~~~~^^
|
||||
File "/root/miniconda3/lib/python3.13/site-packages/cv2/__init__.py", line 153, in bootstrap
|
||||
native_module = importlib.import_module("cv2")
|
||||
File "/root/miniconda3/lib/python3.13/importlib/__init__.py", line 88, in import_module
|
||||
return _bootstrap._gcd_import(name[level:], package, level)
|
||||
~~~~~~~~~~~~~~~~~~~~~~^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^
|
||||
ImportError: libGL.so.1: cannot open shared object file: No such file or directory
|
||||
|
||||
--
|
||||
what should i do
|
||||
|
||||
---
|
||||
|
||||
### 🤖 Cursor Assistant
|
||||
|
||||
This is a common OpenCV error in headless Linux environments (like HPC clusters or Docker containers). The issue is that OpenCV's Python package requires OpenGL libraries, but they're not installed on your system.
|
||||
|
||||
Here are the solutions:
|
||||
|
||||
## Solution 1: Install OpenGL libraries (Recommended for SLURM/HPC)
|
||||
|
||||
For **Ubuntu/Debian**:
|
||||
```bash
|
||||
sudo apt-get update
|
||||
sudo apt-get install -y libgl1-mesa-glx libglib2.0-0
|
||||
```
|
||||
|
||||
For **CentOS/RHEL**:
|
||||
```bash
|
||||
sudo yum install -y mesa-libGL
|
||||
```
|
||||
|
||||
If you don't have sudo access on an HPC cluster, contact your system administrator to install these packages.
|
||||
|
||||
## Solution 2: Use headless OpenCV
|
||||
|
||||
Install the headless version of OpenCV which doesn't require GUI libraries:
|
||||
|
||||
```bash
|
||||
pip uninstall opencv-python opencv-contrib-python opencv-python-headless
|
||||
pip install opencv-python-headless
|
||||
```
|
||||
|
||||
This is often the best solution for server environments where you don't need display functionality.
|
||||
|
||||
## Solution 3: Set environment variable (temporary workaround)
|
||||
|
||||
If you only need basic OpenCV functionality without GUI:
|
||||
|
||||
```bash
|
||||
export LD_PRELOAD=/usr/lib/x86_64-linux-gnu/libffi.so.7
|
||||
```
|
||||
|
||||
Or add this to your SLURM script before the `accelerate launch` command.
|
||||
|
||||
## For your SLURM script specifically:
|
||||
|
||||
I'd recommend adding this to your script before the training command:
|
||||
|
||||
```bash
|
||||
# Add after the variable definitions, before accelerate launch
|
||||
pip install opencv-python-headless --upgrade --force-reinstall
|
||||
|
||||
DS_SKIP_CUDA_CHECK=1 TOKENIZERS_PARALLELISM=false \
|
||||
accelerate launch \
|
||||
...
|
||||
```
|
||||
|
||||
**Best approach**: Use `opencv-python-headless` since you're running a vision-based RL training job on an HPC cluster where GUI libraries typically aren't available or needed.
|
||||
|
||||
---
|
||||
|
||||
|
||||
*Exported from [Cursor View](https://github.com/saharmor/cursor-view)*
|
||||
Reference in New Issue
Block a user