Synthetic Data Sets empower data scientists to train financial AI models faster on more robust, highly-realistic, and PII-free data.
IBM
2023 - 2025
Won “Best AI Solution — Data Insights & Knowledge Management” at the FinTech Futures Banking Tech Awards USA
Scroll ↓
IBM Synthetic Data Sets are a family of virtually-generated, privacy-compliant data sets that catalyze the enterprise adoption of Large Language Models (LLMs), Generative AI, and Agentic AI by enabling faster access to rich data that businesses need for predictive financial solutions. Our team transformed what started as a research idea into a fully-realized infrastructure product from 0→1.
Business impact
↑ Faster time-to-value
Increase in AI software and hardware (Spyre) POCs across IBM Z, Partner Ecosystem, and Global Sales accounts by providing data that sped up the overall AI model lifecycle process, illustrating product value to customers within weeks rather than months.
User impact
↑ Faster data pre-processing
Data scientists went from spending 6+ months to spending less than 2 weeks on data pre-processing because of the pre-labeled and pre-balanced data, some getting started on AI model training immediately after downloading the data sets.
Contribution
Design Lead
I led the design strategy, research, and visual design for end-to-end user experiences spanning the AI model lifecycle with a focus on product knowledge and data pre-processing. I grew a breadth of skills in partner ecosystem revenue opportunities, pricing, versioning, Data-aaS, and agent-based modeling techniques.
Digital twin isometric illustration, IBM Blue Studio
Introduction
Synthetic Data Sets improve model performance across the AI development lifecycle, enabling teams to optimize risk detection in finance with greater accuracy and confidence.
Build
Build an AI model from scratch using synthetic data that has already been pre-processed for enterprise use cases.
Enhance
Enhance an existing model by fine-tuning it with fully-labeled, hyper-realistic, and high quality synthetic data.
Validate
Validate the outcomes and F1 score of an AI model trained on real data and then separately trained on synthetic data.
Data sets curated for enterprise use cases
01
Payment Card Fraud
Accurate fraud detection keeps customers satisfied and loyal while minimizing financial losses.
Simulated credit cards and credit card holders
Detailed transaction histories
Transactions labeled “yes” or “no” for fraud and linked by fraudster ID for pattern tracking across data sets
02
Core Banking and Money Laundering
Provides labeled data, including global and cash transactions unavailable in real banking data.
Simulated transactions for money laundering,
check fraud and automated push payment fraudCaptures fraud scenarios and laundering activities
Labels for types of laundering
Shared account details across data sets
03
Insurance Claims Fraud
Adds synthetic “what-if” scenarios that cover diverse claim types and fraud cases.
Homeowner, policy, claim and disaster event information
Labels for fraudulent claims
Free-text insights for fraud, loan underwriting, and credit scoring
Faster time-to-value in the AI model lifecycle
IBM Synthetic Data Sets simplifies the path from strategy to deployment by eliminating data acquisition bottlenecks, reducing preprocessing, and enabling immediate model training.
How it works
IBM Synthetic Data Sets are built from a virtual world with virtual people buying from virtual businesses transacting through virtual financial institutions who sometimes deal with virtual fraudsters, money-launderers, or natural disasters. This method is called Agent Based Modeling, where complex systems (society) are simulated by representing collections of autonomous, interacting entities called “agents” (people) that follow predefined rules and interact with their environment (transactions, banking, fraud).
The rules of our virtual world are determined by statistical distributions from the publicly-available anonymized data sources: population statistics (i.e. US Census), company 10k data, national weather data (i.e. National Oceanic Atmospheric Administration), and federal reserve data.
My contribution
As the Design Lead for a small team, my responsibilities stretched across research, design, and strategy. This allowed me to better connect the inherently disjointed evolution between customer needs, product direction, and commercialization that emerging technology typically presents upon inception.
Researcher
How did we get here?
Set up a research infrastructure yielding Insights from customer interviews, ecosystem research, and market validation, that shaped customer segmentation, informed product direction, and revealed commercialization opportunities that extended beyond the initial product vision.
Designer
Where are we going?
Translated complex technical needs into clear and actionable business opportunities through mixed design methods like visualizing product narratives, creating user journey maps, and developing launch-ready marketing enablement. By reducing cognitive complexity, the work accelerated customer understanding and internal alignment.
Strategist
What do we build first?
Partnered across product management, engineering, research, sales, and marketing to define partner ecosystem commercialization opportunities and Data-as-a-Service opportunities. My role was to connect technical innovation with market demand, ensuring the product solved the right problem while creating a viable path to launch.
Research-centered strategy
I built the research infrastructure for a 0→1 enterprise product by creating repeatable systems that transformed internal assumptions and ad hoc client conversations into a scalable product discovery practice.
Assumptions and Questions
Before engaging customers, I introduced an assumptions and questions exercise to create a shared source of truth, expose knowledge gaps, and translate broad uncertainty into a prioritized research plan to address product, market, and user needs. Rather than conducting exploratory interviews without direction, every customer engagement was primed to test specific hypotheses.
The critical unknowns that required direct user input were categorized in two key buckets:
PRODUCT USAGE
How realistic is synthetic data compared to real data (i.e. fraud percentage, same or better F1 score)?
What information does IBM Synthetic Data Sets offer that is impossible to get from real data or other artificial data?
COMMERCIALIZATION
Which value propositions and differentiators resonate most with enterprise buyers?
What technical evidence and proof points are required to establish trust and justify adoption?
Engagement funnel
Working with a new enterprise data product meant there was no established customer discovery process. I introduced a structured engagement pipeline that organized research, and enabled the team to validate assumptions, build institutional knowledge, and make roadmap decisions from cumulative customer evidence rather than isolated conversations. The framework has been adopted across the broader portfolio as a standard practice.
01
Prospecting
Shortlisting customers based on personas, diverse geos, AI maturity, machine generation, at-risk to competitors.
02
Recruitment
Communicate product details between users, client reps, and product specialists. Begin scheduling engagements.
03
Engagement
Gathering input from users on product features and adoption through generative and evaluative research.
04
Usage
Independent product usage through POCs, betas, trials, and purchases.
Research metrics
38 Hours of engagement
Across 3 quarters
43 Engagements
1:1 Interviews
Focus groups
Events
271 Observations
70 Positive product feedback
96 Product recommendations
123 As-is AI workflows
54 Pain points with real data
27 Constructive product feedback
28 Unique clients
Financial Services
Insurance
Banking
Technology
Government
Key learnings from customer engagements
Data scientists are responsible for building, enhancing, and validating AI models, while AI Executives are responsible for the overall delivery of customer-centric AI solutions — but using real data introduces access and compliance restraints that make it difficult to develop competitive AI solutions quickly and efficiently for both users.
People are the lifeblood of product
Data Scientist
End User
PAIN POINTS
Restricted data access
Real data is time-consuming and costly to access, taking up to 6 months to retrieve.
Long data pre-processing times
Real data cleaning, labeling, and balancing is complex and use-case specific.
NEEDS
Data realism to validate models
Referential integrity (connected data sets)
Known ground truth and other pre-labeling
OPPORTUNITY
Little/no data pre-processing burdens with and easily accessible trusted data access.
AI Executive
Decision Maker
PAIN POINTS
Data privacy constraints
Real data includes Personally Identifiable Information (PII) that is non-compliant to use for product development.
Limited scope of user behavior
Customers’ real data leaves out user behaviors and market segments from other companies.
NEEDS
Data realism to ensure explainability
Broader user behavior
Models trained on compliant (PII-free) data
OPPORTUNITY
Faster time to value for AI projects with richer, more holistic user data without compromising PII.
As-is journey for different customer segments
Customer value wasn't evenly distributed across the market — while both enterprise customers and ISVs saw value in synthetic data, ISVs faced a fundamentally different constraint: they lacked direct access to the rich, holistic customer data needed to build AI solutions, forcing them to rely on customer logs and system events as proxies for user behavior. This distinction revealed where Synthetic Data Sets delivered the greatest value and reframed the target customer for the initial launch.
INDEPENDENT SOFTWARE VENDORS (ISV) and MANAGED SERVICE PROVIDERS (MSP)
(Target for initial release)
INDIVIDUAL ENTERPRISE CUSTOMERS
(Target for future releases)
Key insight for revenue opportunities
After identifying ISVs and MSPs as the strongest initial market for Synthetic Data Sets, I explored how these insights could extend beyond our primary customer segment. In the spirit of continuous product discovery, I brought customer research back to IBM’s Partner Ecosystem, who similarly expressed a need develop AI-powered products and services enriched with Synthetic Data Sets — ultimately to bridging the needs of (1) ISVs/MSPs who had no holistic data access and (2) individual enterprise customers who had no data diversity.
“The challenge has always been connecting core banking data with other user retail and purchasing behavior to provide an enterprise view of a user’s behavior to Lines of Business.”
Chief Revenue Officer
Customer Experience Management Company
Go-to-market
I leaned into my craft of visual communication and design to create compelling content that made technical product information easy to consume and consistent with IBM’s brand language.
Technical differentiators
Research quickly revealed that customers struggled to distinguish synthetic data from real, anonymized, and artificial data. Rather than presenting technical features in isolation, I reframed customer insights into a comparison framework to clarify our competitive advantage, address common misconceptions, and equip both product and go-to-market teams with a shared narrative.
Marketing content
As a first-of-its-kind product on IBM Infrastructure, I led all visual design efforts for external facing assets across multiple marketing channels, sales enablement, and social media.
Eminence
IBM Synthetic Data Sets won “Best AI Solution — Data Insights and Knowledge Management” at the Banking Tech Awards 2025 hosted by Fintech Futures
Launching a 0→1 enterprise product isn't just about building the right solution, it's about building trust in an entirely new category. Throughout the project, I approached design as a tool for commercialization, helping translate complex technology into a clear product narrative that customers and industry experts could understand and believe.
That strategy extended beyond research and product design to include external validation through industry recognition. Our work was featured across IBM thought leadership, external publications, and competed alongside established global financial institutions. Our small multidisciplinary team demonstrated that thoughtful product strategy, research-driven storytelling, and clear customer positioning could earn credibility for an emerging technology. For me, the award represented more than recognition — it validated that trust can be intentionally designed and that research, product, marketing, and design together can accelerate market adoption.
“This is exactly what we were looking for — it’s the missing piece of the puzzle to jumpstart an end-to-end advanced householding solution for our banking clients.”
OVATION CXM
Banking Customer Experience Orchestration
“IBM Synthetic Data Sets is great because it provides an end-to-end view of a transaction, and we have an appetite to be proactive about pre-trained risk identification by looking at holistic user behavior.”
PROMONTORY NETHERLANDS
Risk Management Consulting
Future releases
To shape the next generation of IBM Synthetic Data Sets, I initiated a new phase of generative research as a part of our continuous discovery process. The outcomes combined customer feedback with outstanding assumptions from our initial discovery work.
A
Additional schemas
RESEARCH OBJECTIVE
What changes do we anticipate customers might need in future versions of our data sets?
INSIGHT
Customers are interested in adding anonymous, user-driven, unstructured data to augment the “perfect” synthetic data created by Agent Based Modeling. Our team will prioritize adding schemas and data that are aligned to our AI business growth strategy (Traditional AI, LLMs, Generative AI, Agentic AI).
B
02 Data-as-a-Service
RESEARCH OBJECTIVE
How might we enable customers to autonomously customize data as they scale their AI needs?
INSIGHT
Customers will benefit from a Data-aaS user experience where they can generate data based on customized parameters, while maintaining the quality offered in IBM Synthetic Data Sets (known ground truth, data realism, referential integrity, rich attributes), through an intuitive interface and easy download.
01
Add net new customer-generated data from in-house seed data
Pros
More personalized data variation outside of IBM SDS master data set
Competitive similarity to data generators
Cons
Potential user expectation to incorporate their real data as seed, compromising the original benefit
More development and design resources needed for prototyping and feedback
02
Subset entire original XL synthetic data sets based on customer needs
Pros
User can download any size of data set
“Build your own” data sets with various schemas relevant to specific use cases
Cons
Less user customization with their own seed data
Unclear how updates to the master data set would affect user generated data over time
Reflection
This project reinforced my belief that design is a powerful tool to reduce uncertainty through interrogation and iteration. I empowered our team achieve less uncertainty, which meant more trust — a concept applied to both how I worked with my team as well as how I served our users.
Alignment and repeatable processes
Trust in our ways of working
Product positioning
Trust that this product is differentiated
User discovery
Trust that we're solving the right problem
Industry recognition
Trust from the market
Commercialization
Trust that customers understand the value
Responsible AI
Trust in the product itself
More information
Eminence
IBM Community Blog
IBM targets mainframe customers with prebuilt AI training modules
Network World
IBM Redbooks
Erik Altman, Dipali Aphale, Joy Deng, Yadu Nandan B, Saurabh Srivastava, Kelly Xiang
Research publications
Arxiv
Erik Altman, Jovan Blanuša, Luc von Niederhäusern, Béni Egressy, Andreea Anghel, Kubilay Atasu
Synthesizing credit card transactions- This link opens in a new tab
Association for Computing Machinery
Erik Altman
FraudGT: A simple, effective, and efficient graph transformer for financial fraud detection
IBM
Lin Junhong, Xiaojie Guo, Yada Zhu, Samuel Mitchell, Erik Altman, Julian Shun
Association for Computing Machinery
Jovan Blanuša, Maximo Cravero Baraja Andreea Anghel, Luc von Niederhäusern, Erik Altman, Haris Pozidis, Kubilay Atasu