停止推遲數據問題
- 一處
- 統一實驗、模擬、文獻
- AI 就緒
- 從一開始就可以進行模型訓練
- 內部部署
- 保留資料主權
- 逆向設計
- 從數據到下一步實驗
研究資料至今仍散落各處
實驗結果分散在試算表、本機硬碟和電子郵件中。論文結論被鎖在 PDF 的表格與圖片裡,同一個物性被三個研究團隊用三種名稱稱呼。
破碎的資料
實驗紀錄依工具、格式和成員分散存放,無法一覽。
重複的實驗
缺少指引下一步的依據,同樣的條件被一遍遍重複。
割裂的分析
資料收集、建模與最佳化各自運行在不同環境中。
鎖在論文裡的知識
已發表的結果留在表格和圖片中,而沒有人有時間去謄錄。
不一致的術語
同一個物性,各研究團隊的叫法不同,單位寫法也各異。
為每個階段提供精準的工具
在收集、前處理、訓練這條三段式流程之外,還並行運行著知識圖譜、逆向設計,以及理解當前頁面情境的助理。
文獻擷取
可透過 PDF、DOI 或 PubMed ID 登錄論文,也可訂閱期刊讓新文章自動匯入。D3Square 讀取全文後,將表格與圖片頁面重新算繪為高解析度影像,使視覺模型能夠讀出文字剖析器悄然弄錯的數值。
- Crossref 與 PubMed 檢索、DOI 匯入、期刊定期擷取
- 擷取摘要、關鍵結論、實體、關係與表格資料
- 表格、圖表與光譜採用 200 DPI 視覺擷取
- 每個數值都帶有來源頁碼,可一鍵核驗
- 將擷取的數值直接對應到資料桶的欄位
本體論(Ontology)
兩個實驗室量測同一個物性,資料卻無法合併——名稱不同,單位不同。本體論概念將這些變體歸併,把單位正規化為 SI,並在超出規格的數值進入訓練集之前將其隔離。
- 每個概念都帶有同義詞、單位、定義與階層
- 單位剖析與 SI 換算,並記錄 QUDT 識別碼
- 新概念與約束條件須經審核與核准
- OWL 推理導出隱含關係,並保留各自的依據
- 支援 TTL 匯入匯出與 EMMO 快照比對,確保互通性
知識圖譜
從文獻中擷取的實體與關係,會匯入一張可以直接檢視的圖譜。依類型篩選、沿關係跳轉、把結果交給助理,或將選取範圍直接送入訓練用資料桶。
- 實體依材料、物性、方法、期刊、作者等類型劃分
- 依類型篩選、隱藏孤立節點、依跳數展開
- 向助理提問,得到基於圖譜的回答
- 將圖譜選取範圍匯出到資料桶用於訓練
- 結果遵循資料權限——無存取權限,也不會出現在依據中
資料桶
上傳實驗資料集後自動完成驗證。缺失值、類別編碼與離群值偵測在同一條前處理流程中完成。
- 上傳時自動驗證並偵測錯誤
- 缺失值與離群值前處理
- 相關性分析與視覺化
- 版本管理與歷史還原
模型訓練
直接從資料桶出發訓練模型,無需撰寫程式碼。以 R²、MAE、RMSE 比較各次執行,再檢視模型究竟學到了什麼——平台會用平實的語言寫出解釋,並附上注意事項。
- 無程式碼訓練:GPR、XGBoost、線性模型、深度學習
- 交叉驗證,並對 R² · MAE · RMSE 並列比較
- SHAP 特徵重要度,標出主導因子
- 平實語言的結果報告:摘要、關鍵結論、注意事項
- 發布到預測分頁後,即可直接用於預測與逆向設計
逆向設計
先設定目標物性,已發布的模型即可反推出滿足條件的成分與製程條件。共五步——匯入模型、對應角色、設定設計變數、搜尋、執行。
- 基於已發布模型執行單目標與多目標最佳化
- 基因演算法、貝氏最佳化、網格與隨機搜尋
- 柏拉圖前緣給出值得製備的候選方案
- 在約束條件範圍內探索設計空間
- 工作階段、執行與結果依專案保存
主動學習
推薦下一步值得量測的樣品。以最少的實驗獲得最大資訊增益,降低研究成本與時程。
- Expected Improvement 與 UCB 採集策略
- 基於 Thompson Sampling 的推薦
- 在迭代循環中不斷精化模型
- 將實驗成本控制在最低
AI 助理
助理知道你目前在哪個頁面。在訓練結果頁會提議比較效能,在知識圖譜頁會基於圖譜作答,在儀表板頁會告訴你有什麼變化。你下達指令,它就直接操作平台。
- 感知頁面情境——為當前頁面提出下一步建議
- 混合檢索:文件向量檢索與圖譜走訪並用
- 執行平台操作:資料桶、訓練、模型、逆向設計、CAE、DOE
- 多語言提問以原文和譯文兩種形式檢索
- 回答僅限於帳號被允許檢視的範圍
不是通用資料工具,而是研發營運平台。
實驗、庫存、文獻、物性與製程條件都在各自的語境中處理,而不是試算表裡一欄欄沒有名字的數字。真正的差別不在於某一項功能,而在於這個閉環——資料、知識與模型彼此強化。
端到端貫通
收集、分析、訓練、預測與逆向設計是一個平台,而不是五個工具拼接而成。
專為材料研發打造
實驗、庫存、文獻、物性與製程條件都是擁有恰當欄位與單位的一等物件。
可解釋的 AI
效能指標、驅動這些指標的變數,以及值得知曉的侷限——都用平實的語言寫出來。
可重複使用的知識與模型
結構化的文獻知識與已訓練模型不再是一次性產物,而成為下一個專案的起點。
研究的每一項輸入,都匯於一處
實驗室設定、庫存、實驗設計與文獻在同一條連貫的工作流中被收集,並匯入團隊可版本管理、共享與訓練的資料桶。
實驗室設定
只需將地點、設備、規程與分析範本標準化一次,此後的每筆紀錄都會繼承這一結構。
庫存管理
追蹤原料、混合物與試樣的數量與履歷。
實驗
將實驗設計為相互連接的節點,再按照該設計記錄實際發生的情況。
文獻
彙集論文、就地閱讀,並把標註的部分轉為結構化資料。
依權限劃分的資料桶
資料只在有意為之的範圍內共享。使用者無權開啟的資料,也會從圖譜查詢與 AI 回答中排除。
帶版本的團隊資產
資料桶保留歷史紀錄,因此用於訓練的資料集日後可以還原並重新核查。
API 與 LIMS 整合
接入你已經在用的實驗室系統,讓資料收集持續進行,無需任何人重新輸入。
From data to a decision
Collection, training and inverse design are four segments of one loop rather than separate projects. The whole design is that each segment leaves its result in the shape the next one needs.
What "AI-ready" actually demands
What blocks model training is usually the shape of the data, not the amount of it. If each researcher names the same material differently, if equipment conditions go unrecorded, if failed runs are never written down, no volume of records will train anything. D3Square has you define materials, equipment, equipment variables and research templates first, so that recording an experiment is already recording a structured row. The disappearance of a cleanup step is a consequence of that, not a feature bolted on afterwards.
Training and validating predictive models
Accumulated data trains property-prediction models, and models are compared side by side on the same screen. What matters is not the point estimate but the uncertainty — knowing which composition ranges the model is unsure about is what tells you where an experiment is worth spending. Which variables drive the outcome surfaces along the way, so the model doubles as a summary of the phenomenon.
Inverse design — from target to composition
The usual calculation takes a composition and returns a property. Inverse design runs the other way: it takes a target property range as the constraint and searches for compositions and process conditions likely to satisfy it. Candidates come back ranked, each carrying how confident the model is about it.
Active learning — choosing the next experiment
When the number of experiments is limited, which one to run next for the most information gained is itself a calculable question. Active learning balances candidates that look good against candidates the model understands least, recommends the next experiment, and closes the loop as that result returns as data.
Why the data stays inside
The LLM used for analysis runs inside your organisation. On an on-premises deployment, the core of your R&D — compositions, process conditions, records of what failed — is never transmitted to an external cloud, and analysis and training proceed with data sovereignty intact.
How it meets the tools you already use
Nothing requires abandoning the spreadsheets, output files and literature notes a team already keeps. Mapping the fields of those existing records onto templates while defining the lab puts past and future data into one workspace with the same structure. Data not leaving with the person who produced it is why university groups use this for handover.
從實驗檯到最佳設計
只需一次設定實驗室,從下一個循環起資料便自行累積,模型自行成長。
-
01
匯入文獻
透過 PDF、DOI 或期刊訂閱匯入論文。內文、表格與圖片都會帶著來源頁碼輸出。
-
02
標準化
將擷取的術語對應到本體論概念,把單位正規化為 SI,並隔離違反規則的數值。
-
03
收集資料
將實驗資料集上傳至資料桶,驗證與前處理自動完成。
-
04
探索與分析
透過相關性分析、散佈圖與分布檢查,理解手中資料的結構。
-
05
火車模型
比較多種演算法,為目標物性挑選最合適的預測模型。
-
06
發布與預測
發布經過驗證的模型,用於預測新實驗條件下的結果。
-
07
最佳化
執行多目標最佳化,找出值得一試的設計參數組合。
平台落地之後的變化
| 領域 | 導入前 | 導入 D3Square 後 |
|---|---|---|
| 資料管理 | 個人電腦上的檔案、試算表,以及彼此不相連的系統。 | 實驗、文獻與分析紀錄匯聚在同一個平台上累積。 |
| 文獻運用 | 洞見隨專案一同流失,也隨人一同流失。 | 由領域本體論結構化的知識圖譜,任何人都能查詢。 |
| 模型訓練 | 專業人員在獨立環境中撰寫程式碼。 | 無程式碼訓練,並給出全團隊都能讀懂的平實解釋。 |
| 候選方案的探索 | 為了找到條件,實驗一遍又一遍地重做。 | 從目標物性出發進行逆向設計,由它給出條件。 |
使用地點
韓國政府研究所
合金成分設計 - 使用 D3Square 進行實驗資料整合和逆向設計工作流程。
韓國企業研發
二次電池材料團隊累積成分和製程數據,並以此為基礎做出決策。
大學(多所)
各種實驗室使用 D3Square 進行實驗室級數據資產建構以及學生之間更順暢的交接。
適用於多種研究領域
- 材料科學
- 化工製程
- 製造品質管理
- 能源研究
D3Square — frequently asked questions
What problem does D3Square solve?
Research data scattered across researchers and projects, and never reused. D3Square collects it in an AI-ready form from the experiment design stage onward, so model training and inverse design continue in the same place.
Can we bring in the experimental data we already have?
Define your materials, equipment, equipment variables and research templates once, and existing records accumulate in the same structure as new ones. Simulation results and literature data are unified into the same workspace.
Does sensitive R&D data leave our organisation?
Analysis runs on an LLM operated inside your organisation, so data is never sent to an external cloud. An on-premises deployment keeps data sovereignty intact.
How is it different from Materials Square?
Materials Square is where simulations run; D3Square is where experimental, computational and literature data accumulate and decide the next experiment. Results become assets in D3Square, and inverse design and active learning propose what to calculate next.
How do we get started?
Tell us the shape your data is in today and what your R&D is aiming at, and we will map out where to start collecting. Begin with a PoC consultation.