新品 材料研究数据平台——收集、建模、设计返回

D3Square

从实验设计的第一天开始就提供人工智能就绪的数据。将分散的计算和实验记录集中到一处,并将其转变为决策资产。

  • 从实验设计开始收集 AI 就绪数据
  • 从流程上自动记录日常实验的虚拟实验室
  • 本地LLM分析——数据无需离开企业内部
  • 逆向设计和主动学习提出了下一个实验

停止推迟数据问题

一处
统一实验、模拟、文献
AI 就绪
从一开始就可以进行模型训练
本地部署
保留数据主权
逆向设计
从数据到下一步实验
挑战

研究数据至今仍散落各处

实验结果分散在电子表格、本地硬盘和邮件中。论文结论被锁在 PDF 的表格与图片里,同一个物性被三个课题组用三种名称称呼。

碎片化的数据

实验记录按工具、格式和成员分散存放,无法一览。

重复的实验

缺少指引下一步的依据,同样的条件被一遍遍重复。

割裂的分析

数据收集、建模与优化各自运行在不同环境中。

锁在论文里的知识

已发表的结果留在表格和图片中,而没有人有时间去誊录。

不一致的术语

同一个物性,各课题组的叫法不同,单位写法也各异。

核心功能

为每个阶段提供精准的工具

在收集、预处理、训练这条三段式流水线之外,还并行运行着知识图谱、逆向设计,以及理解当前页面情境的助手。

文献抽取

可通过 PDF、DOI 或 PubMed ID 登记论文,也可订阅期刊让新文章自动汇入。D3Square 读取全文后,将表格与图片页面重新渲染为高分辨率图像,使视觉模型能够读出文本解析器悄然弄错的数值。

  • Crossref 与 PubMed 检索、DOI 导入、期刊定期抓取
  • 提取摘要、关键结论、实体、关系与表格数据
  • 表格、图表与光谱采用 200 DPI 视觉抽取
  • 每个数值都带有来源页码,可一键核验
  • 将抽取的数值直接映射到数据桶的列

本体(Ontology)

两个实验室测量同一个物性,数据却无法合并——名称不同,单位不同。本体概念将这些变体归并,把单位规范化为 SI,并在超出规格的数值进入训练集之前将其隔离。

  • 每个概念都带有同义词、单位、定义与层级
  • 单位解析与 SI 换算,并记录 QUDT 标识符
  • 新概念与约束条件须经审核与批准
  • OWL 推理导出隐含关系,并保留各自的依据
  • 支持 TTL 导入导出与 EMMO 快照比对,保障互操作性

知识图谱

从文献中提取的实体与关系,会汇入一张可以直接查看的图谱。按类型筛选、沿关系跳转、把结果交给助手,或将选中范围直接送入训练用数据桶。

  • 实体按材料、物性、方法、期刊、作者等类型划分
  • 按类型筛选、隐藏孤立节点、按跳数展开
  • 向助手提问,得到基于图谱的回答
  • 将图谱选中范围导出到数据桶用于训练
  • 结果遵循数据权限——无访问权限,也不会出现在依据中

数据桶

上传实验数据集后自动完成校验。缺失值、类别编码与离群点检测在同一条预处理流水线中完成。

  • 上传时自动校验并检测错误
  • 缺失值与离群点预处理
  • 相关性分析与可视化
  • 版本管理与历史还原

模型训练

直接从数据桶出发训练模型,无需编写代码。以 R²、MAE、RMSE 比较各次运行,再查看模型究竟学到了什么——平台会用平实的语言写出解释,并附上注意事项。

  • 无代码训练:GPR、XGBoost、线性模型、深度学习
  • 交叉验证,并对 R² · MAE · RMSE 并列比较
  • SHAP 特征重要度,标出主导因子
  • 平实语言的结果报告:摘要、关键结论、注意事项
  • 发布到预测标签页后,即可直接用于预测与逆向设计

逆向设计

先设定目标物性,已发布的模型即可反推出满足条件的成分与工艺条件。共五步——导入模型、映射角色、设定设计变量、搜索、执行。

  • 基于已发布模型执行单目标与多目标优化
  • 遗传算法、贝叶斯优化、网格与随机搜索
  • 帕累托前沿给出值得制备的候选方案
  • 在约束条件范围内探索设计空间
  • 会话、运行与结果按项目保存

主动学习

推荐下一步值得测量的样品。以最少的实验获得最大信息增益,降低研究成本与周期。

  • Expected Improvement 与 UCB 采集策略
  • 基于 Thompson Sampling 的推荐
  • 在迭代循环中不断精化模型
  • 将实验成本控制在最低

AI 助手

助手知道你当前在哪个页面。在训练结果页会提议比较性能,在知识图谱页会基于图谱作答,在仪表板页会告诉你有什么变化。你下达指令,它就直接操作平台。

  • 感知页面情境——为当前页面提出下一步建议
  • 混合检索:文档向量检索与图谱遍历并用
  • 执行平台操作:数据桶、训练、模型、逆向设计、CAE、DOE
  • 多语言提问以原文和译文两种形式检索
  • 回答仅限于账户被允许查看的范围
为什么选择 D3Square

不是通用数据工具,而是研发运营平台。

实验、库存、文献、物性与工艺条件都在各自的语境中处理,而不是电子表格里一列列没有名字的数字。真正的差别不在于某一项功能,而在于这个闭环——数据、知识与模型彼此强化。

端到端贯通

收集、分析、训练、预测与逆向设计是一个平台,而不是五个工具拼接而成。

专为材料研发打造

实验、库存、文献、物性与工艺条件都是拥有恰当字段与单位的一等对象。

可解释的 AI

性能指标、驱动这些指标的变量,以及值得知晓的局限——都用平实的语言写出来。

可复用的知识与模型

结构化的文献知识与已训练模型不再是一次性产物,而成为下一个项目的起点。

收集

研究的每一项输入,都汇于一处

实验室配置、库存、实验设计与文献在同一条连贯的工作流中被收集,并汇入团队可版本管理、共享与训练的数据桶。

实验室配置

只需将地点、设备、规程与分析模板标准化一次,此后的每条记录都会继承这一结构。

库存管理

追踪原料、混合物与试样的数量与履历。

实验

将实验设计为相互连接的节点,再按照该设计记录实际发生的情况。

文献

汇集论文、就地阅读,并把标注的部分转为结构化数据。

按权限划分的数据桶

数据只在有意为之的范围内共享。用户无权打开的数据,也会从图谱查询与 AI 回答中排除。

带版本的团队资产

数据桶保留历史记录,因此用于训练的数据集日后可以还原并重新核查。

API 与 LIMS 集成

接入你已经在用的实验室系统,让数据收集持续进行,无需任何人重新录入。

Methodology

From data to a decision

Collection, training and inverse design are four segments of one loop rather than separate projects. The whole design is that each segment leaves its result in the shape the next one needs.

What "AI-ready" actually demands

What blocks model training is usually the shape of the data, not the amount of it. If each researcher names the same material differently, if equipment conditions go unrecorded, if failed runs are never written down, no volume of records will train anything. D3Square has you define materials, equipment, equipment variables and research templates first, so that recording an experiment is already recording a structured row. The disappearance of a cleanup step is a consequence of that, not a feature bolted on afterwards.

Training and validating predictive models

Accumulated data trains property-prediction models, and models are compared side by side on the same screen. What matters is not the point estimate but the uncertainty — knowing which composition ranges the model is unsure about is what tells you where an experiment is worth spending. Which variables drive the outcome surfaces along the way, so the model doubles as a summary of the phenomenon.

Inverse design — from target to composition

The usual calculation takes a composition and returns a property. Inverse design runs the other way: it takes a target property range as the constraint and searches for compositions and process conditions likely to satisfy it. Candidates come back ranked, each carrying how confident the model is about it.

Active learning — choosing the next experiment

When the number of experiments is limited, which one to run next for the most information gained is itself a calculable question. Active learning balances candidates that look good against candidates the model understands least, recommends the next experiment, and closes the loop as that result returns as data.

Why the data stays inside

The LLM used for analysis runs inside your organisation. On an on-premises deployment, the core of your R&D — compositions, process conditions, records of what failed — is never transmitted to an external cloud, and analysis and training proceed with data sovereignty intact.

How it meets the tools you already use

Nothing requires abandoning the spreadsheets, output files and literature notes a team already keeps. Mapping the fields of those existing records onto templates while defining the lab puts past and future data into one workspace with the same structure. Data not leaving with the person who produced it is why university groups use this for handover.

How it works

从实验台到最优设计

只需一次配置实验室,从下一个循环起数据便自行累积,模型自行成长。

  1. 01

    导入文献

    通过 PDF、DOI 或期刊订阅导入论文。正文、表格与图片都会带着来源页码输出。

  2. 02

    标准化

    将抽取的术语对应到本体概念,把单位规范化为 SI,并隔离违反规则的数值。

  3. 03

    收集数据

    将实验数据集上传至数据桶,校验与预处理自动完成。

  4. 04

    探索与分析

    通过相关性分析、散点图与分布检查,理解手中数据的结构。

  5. 05

    火车模型

    比较多种算法,为目标物性挑选最合适的预测模型。

  6. 06

    发布与预测

    发布经过验证的模型,用于预测新实验条件下的结果。

  7. 07

    优化

    运行多目标优化,找出值得一试的设计参数组合。

改变了什么

平台落地之后的变化

领域 导入前 导入 D3Square 后
数据管理 个人电脑上的文件、电子表格,以及彼此不相连的系统。 实验、文献与分析记录汇聚在同一个平台上累积。
文献利用 洞见随项目一同流失,也随人一同流失。 由领域本体结构化的知识图谱,任何人都能查询。
模型训练 专业人员在独立环境中编写代码。 无代码训练,并给出全团队都能读懂的平实解释。
候选方案的探索 为了找到条件,实验一遍又一遍地重做。 从目标物性出发进行逆向设计,由它给出条件。
Use cases

使用地点

韩国政府研究所

合金成分设计 - 使用 D3Square 进行实验数据集成和逆向设计工作流程。

韩国企业研发

二次电池材料团队积累成分和工艺数据,并以此为基础做出决策。

大学(多所)

各种实验室使用 D3Square 进行实验室级数据资产构建以及学生之间更顺畅的交接。

适用领域

适用于多种研究领域

  • 材料科学
  • 化工工艺
  • 制造质量管理
  • 能源研究
FAQ

D3Square — frequently asked questions

What problem does D3Square solve?

Research data scattered across researchers and projects, and never reused. D3Square collects it in an AI-ready form from the experiment design stage onward, so model training and inverse design continue in the same place.

Can we bring in the experimental data we already have?

Define your materials, equipment, equipment variables and research templates once, and existing records accumulate in the same structure as new ones. Simulation results and literature data are unified into the same workspace.

Does sensitive R&D data leave our organisation?

Analysis runs on an LLM operated inside your organisation, so data is never sent to an external cloud. An on-premises deployment keeps data sovereignty intact.

How is it different from Materials Square?

Materials Square is where simulations run; D3Square is where experimental, computational and literature data accumulate and decide the next experiment. Results become assets in D3Square, and inverse design and active learning propose what to calculate next.

How do we get started?

Tell us the shape your data is in today and what your R&D is aiming at, and we will map out where to start collecting. Begin with a PoC consultation.

分散在各处的数据汇集在一起​​就会成为一种资产。

告诉我们您的数据目前存放在哪里以及您的研发目标是什么。我们将规划从哪里开始收集。