[论文解读] From Concept to Manufacturing: Evaluating Vision-Language Models for Engineering Design
本论文系统地评估 GPT-4V 与 LLaVA 1.6 34B 在跨越概念到制造阶段的工程设计任务,并发布基准数据集与 prompts,用于未来对 VLM 的评估。
Engineering design is undergoing a transformative shift with the advent of AI, marking a new era in how we approach product, system, and service planning. Large language models have demonstrated impressive capabilities in enabling this shift. Yet, with text as their only input modality, they cannot leverage the large body of visual artifacts that engineers have used for centuries and are accustomed to. This gap is addressed with the release of multimodal vision-language models (VLMs), such as GPT-4V, enabling AI to impact many more types of tasks. Our work presents a comprehensive evaluation of VLMs across a spectrum of engineering design tasks, categorized into four main areas: Conceptual Design, System-Level and Detailed Design, Manufacturing and Inspection, and Engineering Education Tasks. Specifically in this paper, we assess the capabilities of two VLMs, GPT-4V and LLaVA 1.6 34B, in design tasks such as sketch similarity analysis, CAD generation, topology optimization, manufacturability assessment, and engineering textbook problems. Through this structured evaluation, we not only explore VLMs' proficiency in handling complex design challenges but also identify their limitations in complex engineering design applications. Our research establishes a foundation for future assessments of vision language models. It also contributes a set of benchmark testing datasets, with more than 1000 queries, for ongoing advancements and applications in this field.
研究动机与目标
- 评估视觉-语言模型如何处理结合草图、图纸与文本的多模态工程设计任务。
- 构建标准化的基准与数据集,用以评估工程设计中的 VLM。
- 提供定性与定量分析,识别 VLM 在设计场景中的能力与局限。
- 提供基线评估,指导未来在工程设计领域的 VLM 发展。
提出的方法
- 开发以图像为主要输入、辅以简短文本提示的提示与实验。
- 进行了超过 1000 次查询,以评估 GPT-4V 在包括设计相似性、早期草图描述、CAD 生成、拓扑优化理解、可制造性评估、加工特征识别、缺陷识别、教材题目与空间推理等任务中的表现。
- 使用相同任务和数据集,将 GPT-4V 与开源 VLM LLaVA 1.6 34B 进行对比。
- 提供精确的提示与模型响应,以实现基准任务的可重复性。
实验结果
研究问题
- RQ1视觉-语言模型是否能有效执行同时使用视觉和文本输入的工程设计任务?
- RQ2与人类基线相比,VLM 在概念设计、详细设计、制造/检验以及教育相关任务上的表现如何?
- RQ3在工程设计情境中,VLM 的局限性与失效模式有哪些,基准如何推动未来改进?
- RQ4标准化的数据集与提示能否实现对不同 VLM 在工程设计任务上的公平比较?
主要发现
- GPT-4V 在设计相似性任务(94.0%)中的自洽性高,在 360 组三元组中最小化传递性违规(5),与人类评审者相当或更好。
- GPT-4V 生成的想法图按设计特征逻辑聚类(如奶泡机杯与自行车),与人类生成的图类似。
- 当草图中存在手写文本时,描述匹配任务显示完美准确率(10/10);没有手写文本时,性能下降但仍高于随机,并在去除 “None of the above” 选项时有所提升。
- 通过草图进行描述生成,GPT-4V 能生成与设计内容一致的描述性文本,结果受草图质量影响;定性提示可产生有信息量的描述。
- 该研究提供包含超过 1000 次查询的数据集,并公开输入/提示/答案,以便未来对 VLM 在工程设计中的基准测试。
更好的研究,从现在开始
从阅读论文到最终审阅,大幅缩短您的研究时间。
无需绑定信用卡
本解读由 AI 生成,并经人工编辑审核。