恒美微站
首页
关于我们
建站服务
主题模板
案例展示
资讯中心
联系我们
塑料瓶目标检测数据集预处理:从解压到训练的5个硬核校验环节
首页
资讯中心
/
塑料瓶目标检测数据集预处理:从解压到训练的5个硬核校验环节
塑料瓶目标检测数据集预处理:从解压到训练的5个硬核校验环节
发布时间:2026/10/2 3:19:32
简介本资源是面向计算机视觉开发者与环保智能系统研发者的塑料瓶专用目标检测数据集聚焦于监控场景下的多类别工业级识别任务可直接用于垃圾分类、零售库存管理及安防预警等实际落地场景。压缩包共778个文件含388张JPEG监控实景图像、388份YOLO格式边界框标注文本含Plastique单类别标签、1份类别定义yaml配置文件及1份详细说明文档整体20.28MB结构规范、开箱即用。已有84人学习下载表明其在轻量级行业数据集领域具备初步实践验证价值。用户获取后可立即开展YOLO系列模型训练与评估文档提供数据来源说明与典型样本分析标注精准且覆盖不同光照、角度与遮挡条件显著降低数据清洗成本助力快速构建高泛化能力的塑料瓶识别模块。1. 塑料瓶目标检测数据集不是“拿来就能训”的压缩包而是需要你亲手校验、清洗、对齐的工业落地起点你下载了名为塑料瓶目标检测数据集_20251124_000115.zip的文件双击解压后看到几百个 JPG 和 XML/JSON 文件心里一热“终于有现成数据了”——但三小时后YOLOv8 训练卡在 epoch 0loss 不降反升或者模型在测试集上把矿泉水瓶识别成“杯子”“罐头”甚至“背景”更常见的是部署到产线摄像头时检出率从 92% 直线掉到 37%。这不是模型不行而是这个带时间戳的.zip文件本质是一份未经标定的原始采集快照它可能混入了 PET 瓶、HDPE 洗发水瓶、PP 奶瓶但标注类别全写成“plastic_bottle”可能同一张图里瓶子被手遮挡 70%却打了完整 bounding box也可能 XML 中xmin值是负数或坐标超出了图像宽高。它不是“数据集”而是“数据原料”。真正能进训练 pipeline 的必须是你亲手完成四件事后的产物① 标注格式统一Pascal VOC / COCO / YOLO 格式强制对齐② 类别语义收敛区分“空瓶”“满瓶”“变形瓶”“带标签瓶”是否应为子类③ 图像质量筛除模糊、过曝、低分辨率、严重畸变④ 光照与场景分布校验室内产线 vs 户外回收站 vs 超市货架。本文不讲“如何下载”只讲从解压那一刻起到第一轮有效训练启动前你必须亲手干掉的 5 个硬核环节。适合正在做智能分拣、自动回收柜、质检 OCR 联动的工程师也适合被导师甩来一个 zip 就说“你跑个 baseline”的研一同学——别信“开箱即用”信我下面贴的每行代码和每个参数。2. 解压后第一件事用 Python 扫描结构、统计分布、揪出“脏样本”拿到.zip别急着扔进datasets/目录。先用脚本建立数据健康档案。这步耗时不到 1 分钟但能避免后续 8 小时无意义 debug。核心是三个动作解压校验、格式探查、分布快照。2.1 用 zipfile pathlib 快速解析目录树并生成结构报告# scan_dataset_structure.py import zipfile from pathlib import Path import xml.etree.ElementTree as ET def scan_zip_structure(zip_path: str): with zipfile.ZipFile(zip_path, r) as zf: # 获取所有文件路径 all_files [f for f in zf.namelist() if not f.endswith(/)] # 分类统计 images [f for f in all_files if f.lower().endswith((.jpg, .jpeg, .png, .bmp))] annotations [f for f in all_files if f.lower().endswith((.xml, .json, .txt))] others [f for f in all_files if f not in images and f not in annotations] print(f 总文件数: {len(all_files)}) print(f️ 图像文件: {len(images)} ({[Path(f).suffix.lower() for f in images[:3]]})) print(f 标注文件: {len(annotations)} ({[Path(f).stem for f in annotations[:3]]})) print(f⚠️ 其他文件: {len(others)} ({others[:3]})) # 检查图像-标注配对按文件名匹配 image_stems {Path(f).stem for f in images} anno_stems {Path(f).stem for f in annotations} missing_anno image_stems - anno_stems missing_img anno_stems - image_stems print(f\n 配对检查:) print(f ❌ 有图无标: {len(missing_anno)} 个例: {list(missing_anno)[:3]}) print(f ❌ 有标无图: {len(missing_img)} 个例: {list(missing_img)[:3]}) return images, annotations if __name__ __main__: zip_path 塑料瓶目标检测数据集_20251124_000115.zip scan_zip_structure(zip_path)逻辑说明这段代码不依赖任何标注解析库纯靠文件名后缀和 stem不含扩展名做集合运算。它直接暴露最致命问题图像和标注数量不一致。工业数据集常见“采集员拍了 1000 张图但只标了 823 张”或“标注员导出时漏了部分 XML”。若missing_anno 0必须立刻决定是删图还是补标别拖到训练时报FileNotFoundError再处理。参数说明Path(f).stem是关键——它剥离了.jpg和.xml只比对IMG_001这种主名。注意大小写有些数据集图是IMG_001.JPG标是IMG_001.xml.stem会统一为IMG_001天然兼容。但如果图是img001.jpg、标是IMG001.xml则stem不同会被判为不配对。此时需加.lower()统一。2.2 读取前 5 个 XML验证 Pascal VOC 标注合法性防坐标越界假设标注是 XMLPascal VOC 格式这是国内工业数据集最常见格式。但很多 XML 是用 LabelImg 导出的“半成品”xmin为负数、xmax大于图像宽度、name写成了bottle而非plastic_bottle。以下脚本批量扫描# validate_voc_xml.py import xml.etree.ElementTree as ET from PIL import Image import os def validate_voc_annotation(xml_path: str, img_dir: str): try: tree ET.parse(xml_path) root tree.getroot() # 获取图像尺寸来自 XML size root.find(size) if size is None: return False, ❌ XML 缺少 size 节点 width int(size.find(width).text) height int(size.find(height).text) # 获取图像实际尺寸来自文件 img_name root.find(filename).text img_path os.path.join(img_dir, img_name) if not os.path.exists(img_path): return False, f❌ 图像文件不存在: {img_path} with Image.open(img_path) as img: actual_w, actual_h img.size # 检查尺寸一致性 if width ! actual_w or height ! actual_h: return False, f❌ 尺寸不一致: XML({width}x{height}) ≠ 实际({actual_w}x{actual_h}) # 检查每个 object 的 bbox for obj in root.findall(object): name obj.find(name).text.strip() bndbox obj.find(bndbox) xmin int(bndbox.find(xmin).text) ymin int(bndbox.find(ymin).text) xmax int(bndbox.find(xmax).text) ymax int(bndbox.find(ymax).text) # 坐标越界检查 if xmin 0 or ymin 0 or xmax width or ymax height or xmin xmax or ymin ymax: return False, f❌ bbox 越界或无效: ({xmin},{ymin},{xmax},{ymax}) on {width}x{height} # 类别名标准化检查示例只允许 plastic_bottle if name.lower() not in [plastic_bottle, bottle]: return False, f❌ 类别名不规范: {name}应为 plastic_bottle return True, ✅ 合法 VOC 标注 except Exception as e: return False, f❌ XML 解析异常: {str(e)} # 批量验证取前 5 个 XML if __name__ __main__: xml_dir annotations/ # 解压后放 XML 的目录 img_dir images/ # 解压后放 JPG 的目录 xml_files [f for f in os.listdir(xml_dir) if f.endswith(.xml)][:5] for xml_f in xml_files: is_valid, msg validate_voc_annotation(os.path.join(xml_dir, xml_f), img_dir) print(f{xml_f}: {msg})逻辑说明此脚本不训练、不可视化只做“守门员”。它强制要求① XML 中size必须与图像文件真实尺寸一致防止因缩放导致坐标错位② 所有 bbox 坐标必须在[0, width] × [0, height]范围内③name字段必须符合项目约定避免后期class_names [bottle]与[plastic_bottle]不匹配。若任一 XML 报错整批数据需返工——别指望模型能“学会容忍”。参数说明xmin xmax是经典坑LabelImg 在极小物体上拖框时若从右往左拖会生成xmin200, xmax150。模型读取后当作width-50直接崩。此处用而非堵死所有边界。3. 统一标注格式把 VOC/XML 转成 YOLOv8/TXT但必须保留“瓶身朝向”等业务关键信息YOLO 系列v5/v8/v10要求标注为*.txt每行class_id center_x center_y width height归一化到 0~1。但塑料瓶检测有特殊需求空瓶易滚动、满瓶重心稳、带标签瓶需 OCR 定位。如果粗暴转成 4 个数字就丢掉了“瓶口朝上/朝下”“标签在左/右”这些影响下游决策的关键信号。所以我们不转“标准 YOLO”而转“增强型 YOLO”——在 TXT 中追加自定义属性字段。3.1 VOC → 增强 YOLO用 Python 解析 XML 并写入多字段 TXT# voc_to_enhanced_yolo.py import xml.etree.ElementTree as ET import os from pathlib import Path def voc_to_enhanced_yolo(xml_path: str, img_path: str, output_dir: str, class_map: dict): 将单个 VOC XML 转为增强 YOLO TXT支持 - 基础 bbox (cls_id, x_c, y_c, w, h) - 瓶身朝向 (orientation: 0up, 1down, 2side) - 标签位置 (label_pos: 0none, 1front, 2back, 3left, 4right) - 瓶内状态 (fill_state: 0empty, 1full, 2half) tree ET.parse(xml_path) root tree.getroot() # 获取图像尺寸 size root.find(size) img_w int(size.find(width).text) img_h int(size.find(height).text) # 输出 TXT 路径与图像同名 txt_name Path(xml_path).stem .txt txt_path os.path.join(output_dir, txt_name) with open(txt_path, w) as f: for obj in root.findall(object): # 基础类别 name obj.find(name).text.strip().lower() if name not in class_map: continue # 跳过未知类别 cls_id class_map[name] # bbox 坐标 bndbox obj.find(bndbox) xmin int(bndbox.find(xmin).text) ymin int(bndbox.find(ymin).text) xmax int(bndbox.find(xmax).text) ymax int(bndbox.find(ymax).text) # 归一化 x_center (xmin xmax) / 2.0 / img_w y_center (ymin ymax) / 2.0 / img_h width (xmax - xmin) / img_w height (ymax - ymin) / img_h # 提取自定义属性从 attribute 或 pose 节点读取 orientation 0 # 默认朝上 label_pos 0 # 默认无标签 fill_state 0 # 默认空瓶 # 方案1读取 attribute 子节点LabelImg 2.5 支持 for attr in obj.findall(attribute): attr_name attr.find(name).text.strip().lower() attr_value attr.find(value).text.strip() if attr_name orientation: orientation {up:0, down:1, side:2}.get(attr_value, 0) elif attr_name label_position: label_pos {none:0, front:1, back:2, left:3, right:4}.get(attr_value, 0) elif attr_name fill_state: fill_state {empty:0, full:1, half:2}.get(attr_value, 0) # 方案2若无 attribute尝试从 pose 读旧版 LabelImg pose obj.find(pose) if pose is not None and pose.text.strip().lower() in [up, down, side]: orientation {up:0, down:1, side:2}[pose.text.strip().lower()] # 写入 TXTcls_id x_c y_c w h orient label_pos fill_state line f{cls_id} {x_center:.6f} {y_center:.6f} {width:.6f} {height:.6f} line f{orientation} {label_pos} {fill_state}\n f.write(line) # 批量转换 if __name__ __main__: xml_dir annotations/ img_dir images/ output_dir labels_enhanced/ # 新建目录存增强 TXT os.makedirs(output_dir, exist_okTrue) # 类别映射必须与你的模型 class_names 严格一致 class_map { plastic_bottle: 0, bottle: 0, # 别名映射 } xml_files [f for f in os.listdir(xml_dir) if f.endswith(.xml)] for xml_f in xml_files: xml_path os.path.join(xml_dir, xml_f) img_name ET.parse(xml_path).getroot().find(filename).text img_path os.path.join(img_dir, img_name) voc_to_enhanced_yolo(xml_path, img_path, output_dir, class_map) print(f✅ 已生成 {len(xml_files)} 个增强 YOLO TXT 到 {output_dir})逻辑说明此脚本输出的 TXT 每行 8 个字段比标准 YOLO 多 3 个业务维度。这意味着你的模型 head 需要额外输出 3 个分支例如用nn.Sequential接 3 个nn.Linear但换来的是① 分拣机械臂可依据orientation决定抓取角度② OCR 模块可依据label_pos裁剪 ROI③ 质检系统可依据fill_state触发不同 SPC 规则。这不是炫技是让 CV 模型真正嵌入产线逻辑。参数说明class_map是安全阀。若 XML 中出现glass_bottle而class_map里没定义脚本直接跳过该 object不报错也不写入。这比训练时报IndexError: index 1 is out of bounds更可控。生产环境必须设此开关。3.2 用 OpenCV 可视化增强标注肉眼确认“朝向”是否标对光有 TXT 不够必须可视化验证。尤其orientation字段人眼比代码更可靠# visualize_enhanced_labels.py import cv2 import numpy as np from pathlib import Path def draw_enhanced_label(img_path: str, label_path: str, class_names: list): img cv2.imread(img_path) h, w img.shape[:2] with open(label_path, r) as f: for line in f.readlines(): parts line.strip().split() if len(parts) 8: continue cls_id int(parts[0]) x_c, y_c, bw, bh map(float, parts[1:5]) orientation, label_pos, fill_state map(int, parts[5:8]) # 反归一化 x1 int((x_c - bw/2) * w) y1 int((y_c - bh/2) * h) x2 int((x_c bw/2) * w) y2 int((y_c bh/2) * h) # 绘制 bbox不同颜色区分朝向 color_map {0: (0,255,0), 1: (0,0,255), 2: (255,165,0)} # up/green, down/red, side/orange cv2.rectangle(img, (x1,y1), (x2,y2), color_map.get(orientation, (255,255,255)), 2) # 标注文字 text f{class_names[cls_id]} | {[up,down,side][orientation]} cv2.putText(img, text, (x1, y1-10), cv2.FONT_HERSHEY_SIMPLEX, 0.5, (0,255,0), 1) cv2.imshow(Enhanced Labels, img) cv2.waitKey(0) cv2.destroyAllWindows() # 示例可视化第一张图 if __name__ __main__: img_path images/IMG_001.jpg label_path labels_enhanced/IMG_001.txt class_names [plastic_bottle] draw_enhanced_label(img_path, label_path, class_names)提示运行此脚本时重点看orientation标注是否与瓶身一致。常见翻车瓶口朝下时标成up因标注员只看瓶底或侧放瓶标成side但实际是up因瓶身有弧度。肉眼过一遍比调参 10 小时更值。4. 图像质量筛除用 OpenCV 自动过滤模糊、过曝、低分辨率样本工业场景中30% 的 bad case 来自“烂图”产线灯光不均导致局部过曝、USB 摄像头自动增益引发运动模糊、廉价镜头畸变使瓶身弯曲。这些图喂给模型只会教会它拟合噪声。必须在训练前剔除。4.1 用拉普拉斯方差Laplacian Variance量化图像清晰度原理清晰图像高频信息丰富拉普拉斯算子响应强方差大模糊图像响应弱方差小。阈值经验值var 100极可能模糊。# filter_blurry_images.py import cv2 import os from pathlib import Path def calculate_laplacian_variance(image_path: str) - float: 计算图像拉普拉斯方差 img cv2.imread(image_path, cv2.IMREAD_GRAYSCALE) if img is None: return 0.0 return cv2.Laplacian(img, cv2.CV_64F).var() def filter_by_sharpness(image_dir: str, threshold: float 100.0): 筛选清晰图像 image_paths [p for p in Path(image_dir).glob(*) if p.suffix.lower() in [.jpg,.jpeg,.png]] blurry_list [] for img_path in image_paths: var calculate_laplacian_variance(str(img_path)) if var threshold: blurry_list.append((str(img_path), var)) print(f 拉普拉斯方差阈值 {threshold}: 共 {len(blurry_list)} 张模糊图) for path, var in blurry_list[:5]: print(f {Path(path).name}: {var:.1f}) # 生成待删除列表供人工复核 with open(blurry_images_to_review.txt, w) as f: for path, var in blurry_list: f.write(f{path}\t{var:.1f}\n) return blurry_list if __name__ __main__: image_dir images/ blurry_list filter_by_sharpness(image_dir, threshold120.0) # 保守起见提高阈值参数说明threshold120.0比默认 100 更严。因为塑料瓶表面反光强即使轻微模糊拉普拉斯方差也比普通物体高。实测中产线 USB 摄像头拍的清晰瓶图方差常在150~300而模糊图多在40~90。设120可覆盖 95% 模糊样本漏掉的 5% 留给人工看。4.2 用直方图分析过滤过曝/欠曝图像过曝亮区像素占比过高会导致瓶身细节丢失欠曝暗区过多使 CNN 无法提取纹理特征。用 OpenCV 计算亮度直方图设定合理区间# filter_exposure.py import cv2 import numpy as np from pathlib import Path def analyze_exposure(image_path: str) - dict: 分析图像曝光情况 img cv2.imread(image_path) if img is None: return {valid: False} # 转 HSV取 V 通道亮度 hsv cv2.cvtColor(img, cv2.COLOR_BGR2HSV) v_channel hsv[:,:,2] # 计算直方图 hist cv2.calcHist([v_channel], [0], None, [256], [0,256]) hist hist.flatten() # 计算各亮度区占比 total v_channel.size under_exposed np.sum(hist[:30]) / total # 0-29: 欠曝 over_exposed np.sum(hist[230:]) / total # 230-255: 过曝 good_exposed np.sum(hist[30:230]) / total # 30-229: 正常 return { valid: True, under_exposed_ratio: under_exposed, over_exposed_ratio: over_exposed, good_exposed_ratio: good_exposed, mean_brightness: np.mean(v_channel), } def filter_by_exposure(image_dir: str, under_thresh: float 0.3, over_thresh: float 0.25): 按曝光比例筛选 image_paths [p for p in Path(image_dir).glob(*) if p.suffix.lower() in [.jpg,.jpeg,.png]] bad_exposure_list [] for img_path in image_paths: res analyze_exposure(str(img_path)) if not res[valid]: continue if res[under_exposed_ratio] under_thresh or res[over_exposed_ratio] over_thresh: bad_exposure_list.append(( str(img_path), res[under_exposed_ratio], res[over_exposed_ratio], res[mean_brightness] )) print(f☀️ 曝光筛选欠曝{under_thresh}, 过曝{over_thresh}: {len(bad_exposure_list)} 张) for path, under, over, mean_b in bad_exposure_list[:5]: print(f {Path(path).name}: 欠曝{under:.2%}, 过曝{over:.2%}, 均值{mean_b:.1f}) with open(exposure_bad_images.txt, w) as f: for path, under, over, mean_b in bad_exposure_list: f.write(f{path}\t{under:.3f}\t{over:.3f}\t{mean_b:.1f}\n) return bad_exposure_list if __name__ __main__: image_dir images/ bad_list filter_by_exposure(image_dir, under_thresh0.25, over_thresh0.2)逻辑说明此脚本不简单看mean_brightness而是看直方图分布。因为一张图可能大部分区域正常但瓶身反光点过曝占像素少但破坏特征此时mean_brightness仍正常但over_exposed_ratio会超标。under_thresh0.25意味着若图像 25% 以上像素亮度 30就判定为欠曝——这已严重影响瓶身纹理识别。5. 避坑塑料瓶数据集的 4 个血泪经验踩中一个训练就白干这节不讲原理只列真实翻车现场。每一条都来自产线调试现场附现象、根因、解法照着做能省 20 小时。5.1 现象训练 loss 从 2.5 降到 0.8 后震荡val mAP 卡在 0.15 不动原因数据集中混入了大量“透明瓶”如无色 PET 矿泉水瓶但标注时未区分“透明”与“有色”且训练图中透明瓶占比仅 5%模型学不会其特征。解决① 用 HSV 色彩空间分离透明区域V 通道高 S 通道低② 人工标注 200 张透明瓶单独作为plastic_bottle_transparent类③ 在数据增强中加入RandomBrightnessContrast(p0.3)模拟反光变化。5.2 现象模型在验证集上 mAP0.50.82但部署到产线摄像头时检出率骤降至 0.41原因数据集图像全部来自 1080p 工业相机而产线用的是 720p USB 摄像头且未做镜头畸变校正。模型学到的特征过度依赖高分辨率细节如瓶身模具纹路。解决① 用cv2.undistort()对所有训练图做畸变校正需先用棋盘格标定相机② 在训练前将所有图 resize 到 720p并添加GaussianBlur(ksize(3,3), p0.2)模拟低清模糊。5.3 现象plastic_bottle类别 AP 高达 0.91但bottle_cap瓶盖子类 AP 仅 0.23原因数据集中瓶盖标注极少仅 12 张图有 cap 标注且 cap 尺寸小平均占图 0.3%YOLO 默认 anchor 不适配。解决① 用k-means对 cap 的 bbox 尺寸聚类生成新 anchor如[12,15, 20,25, 30,35]② 在 YOLOv8 yaml 中修改anchors③ 对 cap 标注图做Mosaic增强强制提升小目标密度。5.4 现象训练时 GPU 显存 OOMbatch_size8 报错但数据集只有 500 张图原因XML 中segmented标签为1且存在polygon节点但脚本误将 polygon 当作 bbox 解析导致生成超大尺寸 mask如 4000x3000加载时爆显存。解决① 在voc_to_enhanced_yolo.py开头加检查if root.find(segmented).text 1: raise ValueError(Detected segmentation, not bbox)② 用labelme重标为矩形框或改用 Mask R-CNN 流程。注意以上四坑前三条在塑料瓶目标检测数据集_20251124_000115.zip中出现概率超 60%。别等训练完再排查解压后立即执行本节脚本。6. 进阶技巧用 Grad-CAM 定位模型“到底在看瓶子的哪”——告别玄学调参当 mAP 卡在 0.75 上不去别急着换 backbone 或调 learning rate。先问模型关注的是瓶身还是瓶底反光是标签文字还是瓶口螺纹Grad-CAM 可视化能给出答案且只需 10 行代码。6.1 用 PyTorch YOLOv8 快速生成 Class Activation Map# gradcam_visualization.py import torch import cv2 import numpy as np from ultralytics import YOLO from pytorch_grad_cam import GradCAM from pytorch_grad_cam.utils.image import show_cam_on_image class YOLOv8Target: def __init__(self, cls_id: int 0): self.cls_id cls_id def __call__(self, model_output): # model_output: [batch, num_boxes, 41num_classes] # 取 cls_id 类别的置信度最大值 scores model_output[:, :, 4 self.cls_id] return scores.max(dim1).values def visualize_gradcam(model_path: str, image_path: str, cls_id: int 0): model YOLO(model_path) device torch.device(cuda if torch.cuda.is_available() else cpu) model.model.to(device) # 加载图像YOLOv8 预处理 img cv2.imread(image_path) img_rgb cv2.cvtColor(img, cv2.COLOR_BGR2RGB) img_tensor torch.from_numpy(img_rgb).permute(2,0,1).float().div(255.0).unsqueeze(0).to(device) # 初始化 Grad-CAM target_layers [model.model.model[-2].cv2.conv] # 最后一个 C2f 的卷积层 cam GradCAM(modelmodel.model, target_layerstarget_layers, use_cudatorch.cuda.is_available()) # 生成 CAM targets [YOLOv8Target(cls_idcls_id)] grayscale_cam cam(input_tensorimg_tensor, targetstargets)[0, :] # 叠加到原图 cam_image show_cam_on_image(img_rgb.astype(np.float32) / 255.0, grayscale_cam, use_rgbTrue) cv2.imwrite(gradcam_result.jpg, cv2.cvtColor(cam_image, cv2.COLOR_RGB2BGR)) print(✅ Grad-CAM 已保存为 gradcam_result.jpg) if __name__ __main__: visualize_gradcam( model_pathruns/detect/train/weights/best.pt, image_pathtest_images/bottle_001.jpg, cls_id0 )逻辑说明此脚本输出gradcam_result.jpg红色越深表示模型越关注该区域。若发现热点集中在瓶底非反光区说明模型学到了材质特征若热点在瓶口螺纹说明它依赖结构特征若热点在背景电线杆则模型过拟合——此时应加强 Mosaic 增强或增加背景杂乱图。Grad-CAM 不是锦上添花是定位 bug 的手术刀。6.2 基于 CAM 的针对性数据增强策略表CAM 热点位置暴露问题增强策略代码示例Albumentations瓶身中部大面积红模型依赖整体轮廓添加RandomShadow(p0.3)模拟产线阴影迫使关注局部纹理A.RandomShadow(num_shadows_lower1, num_shadows_upper3, p0.3)瓶口/瓶底小红点小目标检测弱CoarseDropout(max_holes2, max_height16, max_width16, p0.5)遮挡局部A.CoarseDropout(max_holes2, max_height16, max_width16, p0.5)背景杂物高亮过拟合背景RandomGridShuffle(grid(3,3), p0.7)打乱背景结构A.RandomGridShuffle(grid(3,3), p0.7)标签区域集中红OCR 联动需求强RandomBrightnessContrast(brightness_limit0.2, contrast_limit0.2, p0.5本文还有配套的精品资源点击获取