Synthetic Data Sets empower data scientists to train financial AI models faster on more robust, highly-realistic, and PII-free data.

IBM
2023 - 2025

Won “Best AI Solution — Data Insights & Knowledge Management” at the FinTech Futures Banking Tech Awards USA

Scroll ↓

IBM Synthetic Data Sets are a family of virtually-generated, privacy-compliant data sets that catalyze the enterprise adoption of Large Language Models (LLMs), Generative AI, and Agentic AI by enabling faster access to rich data that businesses need for predictive financial solutions. Our team transformed what started as a research idea into a fully-realized infrastructure product from 0→1.

Business impact

↑ Faster time-to-value

Increase in AI software and hardware (Spyre) POCs across IBM Z, Partner Ecosystem, and Global Sales accounts by providing data that sped up the overall AI model lifecycle process, illustrating product value to customers within weeks rather than months.

User impact

↑ Faster data pre-processing

Data scientists went from spending 6+ months to spending less than 2 weeks on data pre-processing because of the pre-labeled and pre-balanced data, some getting started on AI model training immediately after downloading the data sets.

Contribution

Design Lead

I led the design strategy, research, and visual design for end-to-end user experiences spanning the AI model lifecycle with a focus on product knowledge and data pre-processing. I grew a breadth of skills in partner ecosystem revenue opportunities, pricing, versioning, Data-aaS, and agent-based modeling techniques.

Digital twin isometric illustration, IBM Blue Studio

Introduction

Synthetic Data Sets improve model performance across the AI development lifecycle, enabling teams to optimize risk detection in finance with greater accuracy and confidence.

Build

Build an AI model from scratch using synthetic data that has already been pre-processed for enterprise use cases.

Enhance

Enhance an existing model by fine-tuning it with fully-labeled, hyper-realistic, and high quality synthetic data.

Validate

Validate the outcomes and F1 score of an AI model trained on real data and then separately trained on synthetic data.

Data sets curated for enterprise use cases

01

Payment Card Fraud

Accurate fraud detection keeps customers satisfied and loyal while minimizing financial losses.

  • Simulated credit cards and credit card holders 

  • Detailed transaction histories

  • Transactions labeled “yes” or “no” for fraud and linked by fraudster ID for pattern tracking across data sets

02

Core Banking and Money Laundering

Provides labeled data, including global and cash transactions unavailable in real banking data.

  • Simulated transactions for money laundering,
    check fraud and automated push payment fraud

  • Captures fraud scenarios and laundering activities

  • Labels for types of laundering

  • Shared account details across data sets

03

Insurance Claims Fraud

Adds synthetic “what-if” scenarios that cover diverse claim types and fraud cases.

  • Homeowner, policy, claim and disaster event information

  • Labels for fraudulent claims

  • Free-text insights for fraud, loan underwriting, and credit scoring

Faster time-to-value in the AI model lifecycle

IBM Synthetic Data Sets simplifies the path from strategy to deployment by eliminating data acquisition bottlenecks, reducing preprocessing, and enabling immediate model training.

How it works

IBM Synthetic Data Sets are built from a virtual world with virtual people buying from virtual businesses transacting through virtual financial institutions who sometimes deal with virtual fraudsters, money-launderers, or natural disasters. This method is called Agent Based Modeling, where complex systems (society) are simulated by representing collections of autonomous, interacting entities called “agents” (people) that follow predefined rules and interact with their environment (transactions, banking, fraud).

The rules of our virtual world are determined by statistical distributions from the publicly-available anonymized data sources: population statistics (i.e. US Census), company 10k data, national weather data (i.e. National Oceanic Atmospheric Administration), and federal reserve data.

My contribution

As the Design Lead for a small team, my responsibilities stretched across research, design, and strategy. This allowed me to better connect the inherently disjointed evolution between customer needs, product direction, and commercialization that emerging technology typically presents upon inception.

Researcher
How did we get here?

Set up a research infrastructure yielding Insights from customer interviews, ecosystem research, and market validation, that shaped customer segmentation, informed product direction, and revealed commercialization opportunities that extended beyond the initial product vision.

Designer
Where are we going?

Translated complex technical needs into clear and actionable business opportunities through mixed design methods like visualizing product narratives, creating user journey maps, and developing launch-ready marketing enablement. By reducing cognitive complexity, the work accelerated customer understanding and internal alignment.

Strategist
What do we build first?

Partnered across product management, engineering, research, sales, and marketing to define partner ecosystem commercialization opportunities and Data-as-a-Service opportunities. My role was to connect technical innovation with market demand, ensuring the product solved the right problem while creating a viable path to launch.

Research-centered strategy

I built the research infrastructure for a 0→1 enterprise product by creating repeatable systems that transformed internal assumptions and ad hoc client conversations into a scalable product discovery practice.

Assumptions and Questions

Before engaging customers, I introduced an assumptions and questions exercise to create a shared source of truth, expose knowledge gaps, and translate broad uncertainty into a prioritized research plan to address product, market, and user needs. Rather than conducting exploratory interviews without direction, every customer engagement was primed to test specific hypotheses.

The critical unknowns that required direct user input were categorized in two key buckets:

PRODUCT USAGE

  1. How realistic is synthetic data compared to real data (i.e. fraud percentage, same or better F1 score)?

  2. What information does IBM Synthetic Data Sets offer that is impossible to get from real data or other artificial data?

COMMERCIALIZATION

  1. Which value propositions and differentiators resonate most with enterprise buyers?

  2. What technical evidence and proof points are required to establish trust and justify adoption?

Engagement funnel

Working with a new enterprise data product meant there was no established customer discovery process. I introduced a structured engagement pipeline that organized research, and enabled the team to validate assumptions, build institutional knowledge, and make roadmap decisions from cumulative customer evidence rather than isolated conversations. The framework has been adopted across the broader portfolio as a standard practice.

01


Prospecting

Shortlisting customers based on personas, diverse geos, AI maturity, machine generation, at-risk to competitors.

02


Recruitment

Communicate product details between users, client reps, and product specialists. Begin scheduling engagements.

03


Engagement

Gathering input from users on product features and adoption through generative and evaluative research.

04


Usage

Independent product usage through POCs, betas, trials, and purchases.

Research metrics

38 Hours of engagement

Across 3 quarters

43 Engagements

1:1 Interviews
Focus groups
Events

271 Observations

70 Positive product feedback
96 Product recommendations
123 As-is AI workflows
54 Pain points with real data
27 Constructive product feedback

28 Unique clients

Financial Services
Insurance
Banking
Technology
Government

Key learnings from customer engagements

Data scientists are responsible for building, enhancing, and validating AI models, while AI Executives are responsible for the overall delivery of customer-centric AI solutions — but using real data introduces access and compliance restraints that make it difficult to develop competitive AI solutions quickly and efficiently for both users.

People are the lifeblood of product

Data Scientist
End User

PAIN POINTS

  1. Restricted data access

    Real data is time-consuming and costly to access, taking up to 6 months to retrieve.

  2. Long data pre-processing times

    Real data cleaning, labeling, and balancing is complex and use-case specific.

NEEDS

  • Data realism to validate models

  • Referential integrity (connected data sets)

  • Known ground truth and other pre-labeling

OPPORTUNITY

Little/no data pre-processing burdens with and easily accessible trusted data access.

AI Executive
Decision Maker

PAIN POINTS

  1. Data privacy constraints

    Real data includes Personally Identifiable Information (PII) that is non-compliant to use for product development.

  2. Limited scope of user behavior

    Customers’ real data leaves out user behaviors and market segments from other companies.

NEEDS

  • Data realism to ensure explainability

  • Broader user behavior

  • Models trained on compliant (PII-free) data

OPPORTUNITY

Faster time to value for AI projects with richer, more holistic user data without compromising PII.

As-is journey for different customer segments

Customer value wasn't evenly distributed across the market — while both enterprise customers and ISVs saw value in synthetic data, ISVs faced a fundamentally different constraint: they lacked direct access to the rich, holistic customer data needed to build AI solutions, forcing them to rely on customer logs and system events as proxies for user behavior. This distinction revealed where Synthetic Data Sets delivered the greatest value and reframed the target customer for the initial launch.

INDEPENDENT SOFTWARE VENDORS (ISV) and MANAGED SERVICE PROVIDERS (MSP)
(Target for initial release)

INDIVIDUAL ENTERPRISE CUSTOMERS
(Target for future releases)

Key insight for revenue opportunities

After identifying ISVs and MSPs as the strongest initial market for Synthetic Data Sets, I explored how these insights could extend beyond our primary customer segment. In the spirit of continuous product discovery, I brought customer research back to IBM’s Partner Ecosystem, who similarly expressed a need develop AI-powered products and services enriched with Synthetic Data Sets — ultimately to bridging the needs of (1) ISVs/MSPs who had no holistic data access and (2) individual enterprise customers who had no data diversity.

“The challenge has always been connecting core banking data with other user retail and purchasing behavior to provide an enterprise view of a user’s behavior to Lines of Business.”

Chief Revenue Officer
Customer Experience Management Company

Go-to-market

I leaned into my craft of visual communication and design to create compelling content that made technical product information easy to consume and consistent with IBM’s brand language.

Technical differentiators

Research quickly revealed that customers struggled to distinguish synthetic data from real, anonymized, and artificial data. Rather than presenting technical features in isolation, I reframed customer insights into a comparison framework to clarify our competitive advantage, address common misconceptions, and equip both product and go-to-market teams with a shared narrative.

Marketing content

As a first-of-its-kind product on IBM Infrastructure, I led all visual design efforts for external facing assets across multiple marketing channels, sales enablement, and social media.

Eminence

IBM Synthetic Data Sets won “Best AI Solution — Data Insights and Knowledge Management” at the Banking Tech Awards 2025 hosted by Fintech Futures

Launching a 0→1 enterprise product isn't just about building the right solution, it's about building trust in an entirely new category. Throughout the project, I approached design as a tool for commercialization, helping translate complex technology into a clear product narrative that customers and industry experts could understand and believe.

That strategy extended beyond research and product design to include external validation through industry recognition. Our work was featured across IBM thought leadership, external publications, and competed alongside established global financial institutions. Our small multidisciplinary team demonstrated that thoughtful product strategy, research-driven storytelling, and clear customer positioning could earn credibility for an emerging technology. For me, the award represented more than recognition — it validated that trust can be intentionally designed and that research, product, marketing, and design together can accelerate market adoption.

“This is exactly what we were looking for — it’s the missing piece of the puzzle to jumpstart an end-to-end advanced householding solution for our banking clients.”

OVATION CXM
Banking Customer Experience Orchestration

“IBM Synthetic Data Sets is great because it provides an end-to-end view of a transaction, and we have an appetite to be proactive about pre-trained risk identification by looking at holistic user behavior.”

PROMONTORY NETHERLANDS
Risk Management Consulting

Future releases

To shape the next generation of IBM Synthetic Data Sets, I initiated a new phase of generative research as a part of our continuous discovery process. The outcomes combined customer feedback with outstanding assumptions from our initial discovery work.

A

Additional schemas

RESEARCH OBJECTIVE

What changes do we anticipate customers might need in future versions of our data sets?

INSIGHT

Customers are interested in adding anonymous, user-driven, unstructured data to augment the “perfect” synthetic data created by Agent Based Modeling. Our team will prioritize adding schemas and data that are aligned to our AI business growth strategy (Traditional AI, LLMs, Generative AI, Agentic AI).

B

02 Data-as-a-Service

RESEARCH OBJECTIVE

How might we enable customers to autonomously customize data as they scale their AI needs?

INSIGHT

Customers will benefit from a Data-aaS user experience where they can generate data based on customized parameters, while maintaining the quality offered in IBM Synthetic Data Sets (known ground truth, data realism, referential integrity, rich attributes), through an intuitive interface and easy download.

01

Add net new customer-generated data from in-house seed data

Pros

  • More personalized data variation outside of IBM SDS master data set

  • Competitive similarity to data generators

Cons

  • Potential user expectation to incorporate their real data as seed, compromising the original benefit

  • More development and design resources needed for prototyping and feedback

02

Subset entire original XL synthetic data sets based on customer needs

Pros

  • User can download any size of data set

  • “Build your own” data sets with various schemas relevant to specific use cases

Cons

  • Less user customization with their own seed data

  • Unclear how updates to the master data set would affect user generated data over time

Reflection

This project reinforced my belief that design is a powerful tool to reduce uncertainty through interrogation and iteration. I empowered our team achieve less uncertainty, which meant more trust — a concept applied to both how I worked with my team as well as how I served our users.

Alignment and repeatable processes
Trust in our ways of working

Product positioning
Trust that this product is differentiated

User discovery
Trust that we're solving the right problem

Industry recognition
Trust from the market

Commercialization
Trust that customers understand the value

Responsible AI
Trust in the product itself

More information

Research publications

Cornell University: Realistic synthetic financial transactions for anti-money laundering models- This link opens in a new tab

Arxiv
Erik Altman, Jovan Blanuša, Luc von Niederhäusern, Béni Egressy, Andreea Anghel, Kubilay Atasu

Synthesizing credit card transactions- This link opens in a new tab

Association for Computing Machinery
Erik Altman

FraudGT: A simple, effective, and efficient graph transformer for financial fraud detection

IBM
Lin Junhong, Xiaojie Guo, Yada Zhu, Samuel Mitchell, Erik Altman, Julian Shun

Real-time subgraph-based feature extraction for financial crime detection- This link opens in a new tab

Association for Computing Machinery
Jovan Blanuša, Maximo Cravero Baraja Andreea Anghel, Luc von Niederhäusern, Erik Altman, Haris Pozidis, Kubilay Atasu