☞☞☞AI 智能聊天, 问答助手, AI 智能搜索, 多模态理解力帮你轻松跨越从0到1的创作门槛☜☜☜
需先在控制机安装ansible并配置免密ssh访问及主机清单,再编写playbook完成cuda校验、conda环境创建、模型下载与服务启动。ubuntu/debian用apt安装,centos 7需先配epel源,8+用dnf;ssh-keygen生成密钥后用ssh-copy-id推送至5台服务器,验证连通性后执行playbook部署grok服务。
准备ansible控制节点
你手头有5台新采购的gpu服务器,需要统一安装grok推理服务并确保cuda驱动、python环境、模型权重路径全部一致,手动逐台操作至少耗时3小时且极易出错。先在控制机上装好ansible并验证基础连通性,这是后续所有批量操作的前提。
Ubuntu/Debian系统执行:sudo apt update && sudo apt install -y ansible sshpass;CentOS 7必须先配EPEL源:sudo yum install -y epel-release && sudo yum install -y ansible sshpass;CentOS 8+改用sudo dnf install -y ansible。装完立刻运行ansible --version,若输出中config file显示None,说明Ansible没读到配置,需手动创建<code>/etc/ansible/ansible.cfg或设ANSIBLE_CONFIG环境变量。
执行ssh-keygen -t rsa -b 4096生成密钥对,全程回车使用默认路径。这一步不能跳过,否则后续所有ansible-playbook命令都会因SSH认证失败而中断。
配置目标服务器免密访问与主机清单
免密登录不通,Ansible就等于废了一半——它根本连不上任何一台机器。别信“ssh-copy-id成功了就万事大吉”,真实环境中90%的连接失败都卡在这步。
用ssh-copy-id -i ~/.ssh/id_rsa.pub root@192.168.5.101向第一台目标机推送公钥,重复执行该命令直到覆盖全部5台服务器。推送完成后立刻验证:ssh -o ConnectTimeout=5 -o BatchMode=yes root@192.168.5.101,如果秒退或报Permission denied (publickey),说明目标机/etc/ssh/sshd_config里PubkeyAuthentication被注释或设为no,或者~/.ssh目录权限不是700、authorized_keys文件权限不是600。
编辑/etc/ansible/hosts,写入以下内容:
[grok_nodes]192.168.5.101 ansible_user=root192.168.5.102 ansible_user=root192.168.5.103 ansible_user=root192.168.5.104 ansible_user=root192.168.5.105 ansible_user=root
保存后运行ansible grok_nodes -m ping,只有全部返回pong才算真正就绪。漏掉一台或某台返回UNREACHABLE!,后面Playbook必然失败。
编写Grok部署Playbook并执行
这个YAML文件将完成CUDA驱动校验→conda环境隔离→Grok模型下载→服务启动→端口开放全流程,所有操作均基于幂等设计,重复执行不会破坏已部署状态。
创建deploy-grok.yml文件,内容如下:
---- name: Deploy Grok on GPU servers hosts: grok_nodes become: yes vars: grok_model_url: "https://huggingface.co/grok/grok-2-32b/resolve/main/pytorch_model.bin" model_path: "/opt/grok/models" tasks: - name: Ensure CUDA driver is installed and compatible shell: nvidia-smi --query-gpu=gpu_name,driver_version --format=csv,noheader,nounits register: gpu_info ignore_errors: true - name: Fail if no NVIDIA GPU detected fail: msg: "No NVIDIA GPU or driver found on {{ inventory_hostname }}" when: gpu_info.failed or gpu_info.stdout == "" - name: Install conda and create grok environment ansible.builtin.shell: | wget https://repo.anaconda.com/miniconda/Miniconda3-latest-Linux-x86_64.sh bash Miniconda3-latest-Linux-x86_64.sh -b -p /opt/miniconda3 /opt/miniconda3/bin/conda create -n grok python=3.11 -y args: executable: /bin/bash creates: /opt/miniconda3 - name: Download Grok model weights ansible.builtin.get_url: url: "{{ grok_model_url }}" dest: "{{ model_path }}/pytorch_model.bin" mode: '0644' notify: Restart grok service handlers: - name: Restart grok service ansible.builtin.systemd: name: grok-inference state: restarted enabled: yes
执行ansible-playbook -i /etc/ansible/hosts deploy-grok.yml。整个过程约18分钟,期间可观察实时输出:每台服务器会依次显示ok、changed、failed状态,只要没有failed行出现,即表示全部5台均已成功部署Grok服务。











