AI is often thought of as the last step in digital transformation. In the domains of materials and chemicals, it should be the inverse.

Image Credit: pixadot.studio/Shutterstock.com
Many people have heard: “You need lots of data before you can do AI.” This is now accepted wisdom in MBA programs, management playbooks, and R&D teams; in materials and chemicals, this is the wrong approach.
The Reality: Data Is Costly and Scarce
In multiple industries, there is a lot of data. In materials and chemicals, however, there is not.
Every data point can cost hundreds or thousands of dollars to collect; experiments take time; equipment is limited; and when working at the technological forefront, the data needed often doesn’t even exist.
A paradox is thus created: one is told they need large datasets to use AI, but creating those datasets is slow, expensive, and uncertain.
The end result is that teams put AI adoption on the back burner while waiting to acquire data that is “AI-ready”. In practice, this often means waiting for undefined periods of time.
The Default Response: Data First, Value Later
When faced with this challenge, many organizations launch large-scale data initiatives.
They hope to:
- Collect historical data
- Standardize units and definitions
- Build an infrastructure that is centralized
The intentions behind these efforts are good; however, they tend to fall into a familiar pattern:
- Months spent discussing data definitions
- Years sunk building systems
- Large investment in infrastructure
Ultimately, there’s no guarantee that:
- The data is relevant to the issues one wants to solve
- The structure supports AI workflows
- The effort produces measurable value
In some cases, teams end up with well-organized data that doesn’t change anything.
A Better Approach: Start with AI
There is a better way to tackle this issue. Rather than waiting for perfect data, teams can start with AI built for the reality of materials R&D:
- Operates with small, imperfect datasets
- Handles missing, noisy, and inconsistent data
- Automatically harmonizes and normalizes data behind the scenes
Crucially, contemporary AI helps you decide what data to produce next.

Image Credit: Citrine Informatics
Let the Model Tell You What Data it Needs
In conventional R&D, experimentation often follows intuition and trial and error; with AI, that doesn’t have to be the case. The model explains what data is required.
The platform can:
- Recommend experiments that will lessen that uncertainty in the most efficient way
- Recognize which experiments are most likely to produce good results
- Underline where the model is uncertain
Rather than building a large dataset upfront, one should:
- Begin with what they have
- Use AI to guide the next best tests
- Produce only the data that actually enhances outcomes
This is a significantly more efficient way to work.

Image Credit: Citrine Informatics
Prove Value First
With the correct tools in place, teams can run projects that are both focused and impactful:
- Maximize a formulation
- Enhance a key property
- Reduce experimental cycles
These projects are about learning at a faster rate. By prioritizing the most informative tests, teams:
- Lessen wasted lab work
- Reach target performance more quickly
- Build helpful data as a byproduct of progress
Let Use Cases Shape Data Strategy
When AI is built in early on, something important happens: companies start to understand their data through the lens of real-life applications.
Rather than asking: “What data should we collect?” teams start asking: “What data actually drives outcomes?”
This shift is extremely important. It ensures that one's data strategy is grounded in real decisions, workflows, and impact, rather than assumptions or generic best practices.
Driving Adoption from the Inside Out
There’s another key advantage that’s often ignored: people buy in. When scientists and engineers see that AI delivers results:
- They understand why structured data is important
- They prioritize enhancing data quality
- They contribute to developing datasets that are reusable
Data quality becomes better not because someone necessarily asked, but because it is obviously valuable.
AI as the Beginning, Not the End
AI is often thought of as the last step in one’s digital transformation journey. In materials and chemicals, the inverse is usually true. AI is not the end of the journey; it’s just the beginning.
Start with tools that can work with the available data; use them to produce value; let that value guide how data changes over time.
Don’t Put the Cart Before the Horse
The idea that one needs perfect, large-scale datasets before using AI is impractical and counterproductive. Instead:
- Start small
- Prove value
- Build momentum
- Let data follow
Do this because in the world of materials product development, the most efficient way to better data is to simply begin using it.

This information has been sourced, reviewed, and adapted from materials provided by Citrine Informatics.
For more information on this source, please visit Citrine Informatics.