恒美微站
首页
关于我们
建站服务
主题模板
案例展示
资讯中心
联系我们
Prometheus Node Exporter Helm Chart 部署指南:基于 charts 仓库的 Kubernetes 节点监控实践
首页
资讯中心
/
Prometheus Node Exporter Helm Chart 部署指南:基于 charts 仓库的 Kubernetes 节点监控实践
Prometheus Node Exporter Helm Chart 部署指南:基于 charts 仓库的 Kubernetes 节点监控实践
发布时间:2026/10/8 7:51:28
【免费下载链接】charts⚠️(OBSOLETE) Curated applications for Kubernetes项目地址https://gitcode.com/gh_mirrors/chart/charts点击查看免费下载导读本文以 charts 仓库中的 stable/prometheus-node-exporter/README.md 为核心文档系统讲解如何在 Kubernetes 集群上通过 Helm 部署 Prometheus Node Exporter覆盖安装/卸载、全部可配置参数、参数默认值以及 DaemonSet、Service、ServiceMonitor、RBAC 等底层模板实现。读完本文你将掌握用一条helm install命令拉起全节点指标采集、按需调整镜像与端口、打通 Prometheus 服务发现并为外部已部署的 Node Exporter 补充 Endpoints 的完整实战方案。一、背景为什么用 Node Exporter 监控 Kubernetes 节点Prometheus Node Exporter 是 Prometheus 生态中最常用的主机指标采集器负责暴露节点的 CPU、内存、磁盘、网络、文件系统等操作系统级指标。在 Kubernetes 中节点层面的指标如node_cpu_seconds_total、node_memory_*、node_filesystem_*都来自 Node Exporter是集群可观测性建设的基础组件。本仓库中的prometheus-node-exporter是一个 Helm 图表chart其元数据定义在 Chart.yamlname: prometheus-node-exporterversion: 1.11.2appVersion: 1.0.1即默认打包的 Node Exporter 镜像版本为 v1.0.1。需要特别说明的是该图表在仓库中已被标记为deprecated: trueREADME 首行即声明 DEPRECATED and moved to https://github.com/prometheus-community/helm-charts——即该项目已停止维护并迁移到 prometheus-community 的 Helm Charts 仓库。对于新部署建议优先使用迁移后的维护版本本文仍以本仓库中的这份图表为准讲解其配置与实现逻辑供在既有环境中沿用此版本的用户参考。二、快速开始安装与卸载TL;DR最简安装$ helm install stable/prometheus-node-exporter该命令会以默认配置在集群中部署 Node Exporter默认每个节点运行一个 PodDaemonSet 形态并创建对应的 Service、ServiceAccount 与 RBAC 资源。指定 release 名称安装$ helm install --name my-release stable/prometheus-node-exporter安装完成后release 名为my-release集群中生成的资源名称遵循my-release-prometheus-node-exporter的命名规则由 _helpers.tpl 中的fullname模板生成超过 63 字符会被截断以符合 DNS 规范。若 release 名已包含 chart 名则直接使用 release 名避免重复拼接。卸载$ helm delete my-release该命令会删除该 release 关联的所有 Kubernetes 组件DaemonSet、Service、ServiceAccount、PSP、ServiceMonitor 等并删除 release 记录。三、核心设计DaemonSet hostNetwork 的逐节点采集方案Node Exporter 采集的是节点操作系统指标因此天然需要“每个节点运行一个实例”。本图表的默认工作负载类型是DaemonSet其定义见 daemonset.yaml关键设计如下hostNetwork: true默认值Pod 直接使用宿主机网络栈配合--web.listen-address$(HOST_IP):9100监听地址使指标可在节点 IP 的 9100 端口直接访问hostPID: truePod 共享宿主机 PID 命名空间配合挂载的/host/proc读取宿主机全部进程信息hostPath 卷挂载将宿主机的/proc与/sys分别只读挂载到容器内/host/proc与/host/sys并通过默认启动参数--path.procfs/host/proc、--path.sysfs/host/sys指向它们从而采集节点级 CPU、内存、磁盘与网络指标HOST_IP 环境变量当service.listenOnAllInterfaces为true默认时HOST_IP 固定为0.0.0.0监听所有接口否则从status.hostIP字段读取 Pod 被分配到的节点 IP只监听该地址见 daemonset.yaml探针容器配置了基于 HTTPGET /的 livenessProbe 与 readinessProbe探针端口取service.port默认 9100便于滚动更新时健康检查updateStrategy默认采用RollingUpdatemaxUnavailable: 1即滚动升级时最多允许 1 个节点上的副本不可用保证采集能力平滑过渡。四、完整参数表与默认值以下为 README 中给出的全部可配置参数默认值均可在 values.yaml 中对应确认ParameterDescriptionDefaultimage.repositoryImage repositoryquay.io/prometheus/node-exporterimage.tagImage tagv1.0.1image.pullPolicyImage pull policyIfNotPresentextraArgsAdditional container arguments[]extraHostVolumeMountsAdditional host volume mounts[]podAnnotationsAnnotations to be added to node exporter pods{}podLabelsAdditional labels to be added to pods{}rbac.createIf true, create use RBAC resourcestruerbac.pspEnabledSpecifies whether a PodSecurityPolicy should be created.trueresourcesCPU/Memory resource requests/limits{}service.typeService typeClusterIPservice.portThe service port9100service.targetPortThe target port of the container9100service.nodePortThe node port of the serviceservice.listenOnAllInterfacesIf true, listen on all interfaces using IP0.0.0.0. Else listen on the IP address pod has been assigned by Kubernetes.trueservice.annotationsKubernetes service annotations{prometheus.io/scrape: true}serviceAccount.createSpecifies whether a service account should be created.trueserviceAccount.nameService account to be used. If not set andserviceAccount.createistrue, a name is generated using the fullname templateserviceAccount.imagePullSecretsSpecify image pull secrets[]securityContextSecurityContextSee values.yamlaffinityA group of affinity scheduling rules for pod assignment{}nodeSelectorNode labels for pod assignment{}tolerationsList of node taints to tolerate- effect: NoSchedule operator: ExistspriorityClassNameName of Priority Class to assign podsnilendpointslist of addresses that have node exporter deployed outside of the cluster[]hostNetworkWhether to expose the service to the host networktrueprometheus.monitor.enabledSet this totrueto create ServiceMonitor for Prometheus operatorfalseprometheus.monitor.additionalLabelsAdditional labels that can be used so ServiceMonitor will be discovered by Prometheus{}prometheus.monitor.namespacenamespace where servicemonitor resource should be createdthe same namespace as prometheus node exporterprometheus.monitor.relabelingsRelabelings that should be applied on the ServerMonitor{}prometheus.monitor.scrapeTimeoutTimeout after which the scrape is ended10sconfigmapsAllow mounting additional configmaps.[]namespaceOverrideOverride the deployment namespace(Release.Namespace)updateStrategyConfigure a custom update strategy for the daemonsetRolling update with 1 max unavailablesidecarsAdditional containers for export metrics to text file[]sidecarVolumeMountVolume for sidecar containers[]五、关键参数深入解读1. 镜像与资源image / resourcesimage: repository: quay.io/prometheus/node-exporter tag: v1.0.1 pullPolicy: IfNotPresent模板中镜像渲染为{{ .Values.image.repository }}:{{ .Values.image.tag }}见 daemonset.yaml即默认quay.io/prometheus/node-exporter:v1.0.1。若需使用镜像仓库镜像可替换image.repository。resources默认留空{}。values.yaml 中给出了推荐示例limits: {cpu: 200m, memory: 50Mi}、requests: {cpu: 100m, memory: 30Mi}。按 charts 社区惯例默认不设资源上限以兼容 Minikube 等小资源环境生产环境建议显式配置。2. Service 暴露方式service.*service.yaml 生成的 Service 默认类型为ClusterIP端口port与targetPort均为 9100端口名metrics。要点默认注解prometheus.io/scrape: true供 Prometheus 基于注解的服务发现直接抓取当service.type为NodePort且设置了service.nodePort时模板会为其追加nodePort字段见 service.yamlservice.targetPort与service.port可分别调整仓库的 CI 用例 ci/port-values.yaml 演示了将两者同时改为 9102 的场景用于验证自定义端口下模板仍可正常渲染。3. 附加启动参数extraArgsNode Exporter 的默认启动参数由模板固定注入--path.procfs/host/proc、--path.sysfs/host/sys、--web.listen-address$(HOST_IP):service.port。需要额外控制采集器时通过extraArgs追加values.yaml 提供了两个高频示例extraArgs: - --collector.diskstats.ignored-devices^(ram|loop|fd|(h|s|v)d[a-z]|nvme\dn\dp)\d$ - --collector.textfile.directory/run/prometheus第一个示例忽略 ram、loop、fd 及 sd/hd/vd/nvme 分区等磁盘设备减少无关指标噪声第二个示例启用 textfile collector从/run/prometheus目录读取自定义文本指标文件通常由 cron 等任务写出的业务自定义指标是扩展 Node Exporter 指标能力的最常见做法。4. 主机目录挂载与 ConfigMapextraHostVolumeMounts / configmapsextraHostVolumeMounts: - name: mountName hostPath: hostPath mountPath: mountPath readOnly: true|false mountPropagation: None|HostToContainer|BidirectionalextraHostVolumeMounts用于把宿主机额外目录如/run/prometheus文本指标目录挂载进容器。模板会同时生成容器volumeMounts与对应的hostPath卷见 daemonset.yaml并支持mountPropagation传播模式。configmaps则用于挂载 ConfigMap 文件configmaps: - name: configMapName mountPath: mountPath5. 边车容器与文本指标卷sidecars / sidecarVolumeMount这是一个值得展开的高级特性借助textfile collector可以在 Node Exporter 旁挂边车sidecar容器如 GPU 厂商的nvidia/dcgm-exporter由边车把指标写入共享的文本文件目录再由 Node Exporter 的 textfile collector 统一暴露。配置方式sidecars: - name: nvidia-dcgm-exporter image: nvidia/dcgm-exporter:1.4.3 sidecarVolumeMount: - name: collector-textfiles mountPath: /run/prometheus readOnly: false模板实现上sidecar 卷以emptyDirmedium: Memory内存盘挂载node-exporter 主容器将该卷挂到同一路径只读边车容器则读写共享见 daemonset.yaml。将--collector.textfile.directory/run/prometheus加入extraArgs后即可打通整条链路。6. 外部节点采集endpoints当部分节点如裸机服务器、虚拟机无法运行 Kubernetes Pod但已经单独部署了 Node Exporter 时可通过endpoints参数把这些外部地址纳入同一 Service 的 Endpointsendpoints: - 10.0.0.1 - 10.0.0.2endpoints.yaml 会生成同名的 Endpoints 资源将列出的 IP 全部绑定到metrics端口固定 9100TCP使 Prometheus 通过该 Service 即可同时抓取集群内 DaemonSet Pod 与集群外 Node Exporter。注意启用endpoints时集群内 Pod 的 Endpoints 由 DaemonSet 自动维护二者以同名 Endpoints 共存。7. 安全上下文与 RBAC / PSPsecurityContextvalues.yaml 默认fsGroup: 65534、runAsGroup: 65534、runAsNonRoot: true、runAsUser: 65534即以非 root 用户nobodyUID 65534运行符合最小权限原则rbac.create / serviceAccount.create均为true模板会创建专用 ServiceAccount名称为fullname模板生成可用serviceAccount.name覆盖imagePullSecrets可注入私有镜像拉取凭据见 serviceaccount.yaml并在 DaemonSet 中引用rbac.pspEnabled默认true当rbac.create为真时会创建policy/v1beta1的 PodSecurityPolicy允许hostPath卷、hostNetwork: true、hostPID: true及 0–65535 的 hostPort见 psp.yaml并配套创建 ClusterRole 与 ClusterRoleBinding 授予该 ServiceAccount 使用权限见 psp-clusterrole.yaml 与 psp-clusterrolebinding.yaml。PSP 仅对启用 PSP 准入控制的旧集群生效在较新集群中该能力已被 Pod Security Admission 取代可视集群版本将rbac.pspEnabled设为false。8. 调度与更新策略affinity / nodeSelector / tolerations / updateStrategytolerations默认- effect: NoSchedule, operator: Exists使 Node Exporter 能调度到所有节点包括被打上污点的节点这是“每节点一个实例”的兜底保证nodeSelector可限定在特定架构/系统节点运行如beta.kubernetes.io/arch: amd64、beta.kubernetes.io/os: linux混合架构集群常用affinity支持节点亲和等高级调度规则values.yaml 附有nodeAffinity注释示例priorityClassName为空时不指定可设置为高优先级保证采集任务不被驱逐updateStrategy默认RollingUpdate / maxUnavailable: 1模板在 DaemonSet 上原样渲染见 daemonset.yaml。9. 命名空间覆盖namespaceOverridenamespaceOverride默认空此时所有资源部署在Release.Namespace。设置后模板中所有资源统一使用覆盖后的命名空间见 _helpers.tpl 中的namespace定义便于在多命名空间组合部署combined charts场景下使用。六、指定参数的方式--set 与 values 文件方式一--set 逐项指定$ helm install --name my-release \ --set serviceAccount.namenode-exporter \ stable/prometheus-node-exporter--set支持keyvalue[,keyvalue]的多项逗号分隔语法适合少量覆盖。方式二-f 指定 YAML 文件$ helm install --name my-release -f values.yaml stable/prometheus-node-exporter适合将完整的自定义配置如上面的 sidecar、extraArgs、endpoints 等组合写入values.yaml后统一注入。两种方式可混用--set优先级更高。七、与 Prometheus 集成注解发现与 ServiceMonitor图表提供了两条开箱即用的抓取路径注解自动发现Service 默认带注解prometheus.io/scrape: true配合 Prometheus 的kubernetes_sd_configsrole: service/endpoints 注解过滤即可自动纳入抓取无需额外配置ServiceMonitorPrometheus Operator 模式设置prometheus.monitor.enabled: true后monitor.yaml 会生成monitoring.coreos.com/v1的 ServiceMonitor按app与release标签选择 Service 的metrics端口并支持additionalLabels追加标签便于 Operator 的 ServiceMonitorSelector 发现scrapeTimeout抓取超时默认10srelabelings自定义 relabel 规则如过滤、重命名、追加指标标签prometheus.monitor.namespace默认与 Node Exporter 同命名空间可指定 ServiceMonitor 所在命名空间。八、验证与访问指标安装完成后可按 Service 类型用 NOTES.txt见 NOTES.txt中的方式访问ClusterIP默认端口转发后访问本机 9100$ kubectl port-forward --namespace ns pod-name 9100NodePort$ export NODE_PORT$(kubectl get --namespace ns -o jsonpath{.spec.ports[0].nodePort} services release-prometheus-node-exporter) $ export NODE_IP$(kubectl get nodes --namespace ns -o jsonpath{.items[0].status.addresses[0].address}) $ echo http://$NODE_IP:$NODE_PORTLoadBalancer等待外部 IP 就绪后访问http://$SERVICE_IP:port。访问后可用curl http://endpoint:9100/metrics验证指标输出如node_cpu_seconds_total、node_memory_*、node_filesystem_*等配合node_uname_info可确认节点身份。九、注意事项与适用前提弃用状态本仓库中的该图表已标记 deprecatedChart.yaml中deprecated: true并迁移至 prometheus-community/helm-charts新环境建议使用迁移后的维护版本本文参数与实现逻辑可供对照迁移hostNetwork 默认开启Pod 复用宿主机网络9100 端口暴露于节点网络请结合网络策略/防火墙评估暴露面如不希望如此可设置hostNetwork: false并调整listenOnAllInterfacesPSP 兼容性默认生成的 PodSecurityPolicy 属于policy/v1beta1仅在集群启用了 PSP 准入时生效新版集群请关闭rbac.pspEnabled自定义端口若通过service.port/service.targetPort更改端口参考 ci/port-values.yaml请同步调整 Prometheus 的抓取配置勿遗漏--web.listen-address中的端口联动关系。十、参考资源核心文档stable/prometheus-node-exporter/README.md默认配置values.yamlChart 元数据Chart.yaml工作负载与网络模板daemonset.yaml、service.yaml、endpoints.yamlPrometheus 集成monitor.yaml安全与 RBACpsp.yaml、serviceaccount.yamlCI 校验示例ci/port-values.yaml命名与命名空间模板templates/_helpers.tpl赞分享【免费下载链接】charts⚠️(OBSOLETE) Curated applications for Kubernetes项目地址https://gitcode.com/gh_mirrors/chart/charts点击查看免费下载相关推荐终极指南如何使用werf实现Kubernetes应用的自动化测试全流程终极指南如何使用werf实现Kubernetes应用的自动化测试全流程 werf是一个强大的解决方案旨在为Kubernetes实现高效且一致的软件交付流程人工智能AI 应用语音移动开发后端桌面应用智能硬件MCP 服务Alluxio Kubernetes 集群监控部署指南基于 Prometheus Grafana 的 Monitor Helm Chart 实践Alluxio Kubernetes 集群监控部署指南基于 Prometheus Grafana 的 Monitor Helm Chart 实践 导读 本存储分布式文件系统缓存大数据基于 helm/charts 仓库的 ChartMuseum 部署实战在 Kubernetes 上自建私有 Helm Chart 仓库基于 helm/charts 仓库的 ChartMuseum 部署实战在 Kubernetes 上自建私有 Helm Chart 仓库 ChartMuseum上一篇RamaLama命令行自动补全提升开发效率的小技巧下一篇如何破解Charles代理工具3分钟完成4.2.7版本激活创作声明:本文部分内容由AI辅助生成(AIGC),仅供参考