← Franklin Dickinson
Cornell Tech · Fall 2025 · Deep Learning

Best practices have preconditions

A multi-architecture experimental evaluation of model selection on credit card fraud detection

A controlled evaluation of XGBoost, FT-Transformer, and BERT architectures on highly imbalanced credit card fraud detection, testing when standard ML practices hold and when they break.

TL;DR
1
Feature type, not model choice, determines which architecture wins.
BERT won on semantic features. XGBoost won on numerical. The ranking reversed completely.
I tested XGBoost, FT-Transformer, and BERT across three fraud detection datasets, five data scales, and five random seeds. BERT dominated on semantic features (0.803 PR-AUC). XGBoost dominated on numerical features (0.865). Which model performed best reversed entirely depending on the feature representation.
2
Cross-domain transfer: shared task label is not a reliable proxy for compatibility.
Two datasets both labeled "fraud detection" had a 20x mismatch in the signal the model needed to learn.
BERT was pre-trained on US e-commerce fraud data (3.5% fraud rate, 434 raw features) then fine-tuned on European card-present data (0.17% fraud rate, 30 PCA features). Both are fraud detection, but the prevalence, feature space, and transaction domain are structurally different. At small data scales, the pre-trained representations provided useful signal. At full data, the fine-tuning gradient overwrote them unstably, and performance collapsed to near-zero. A single-scale evaluation would have called this either a success (at 10%) or a failure (at 100%) and missed the full picture.
3
Standard evaluation practices hid critical failures.
Single-seed runs, naive sampling, and unchecked data scaling each produced false confidence.
Multi-seed validation caught a 4% catastrophic failure rate invisible to single runs. One dataset peaked at 25% with no gain from 75% more data. Stratified downsampling destroyed minority-class diversity, producing a 50% val-to-test gap.
4
Takeaway: Start with the problem, not the model.
A six-step workflow for checking preconditions before committing to an approach.
1
Data characteristics first. Class distribution, feature types, volume, target diversity.
2
Match architecture to features. BERT for semantic. XGBoost for numerical. Fit the model to the data.
3
Test cross-domain transfer. Shared task labels don't guarantee structural compatibility. Run the comparison at multiple scales.
4
Multi-seed from day one. A single score is an anecdote, not evidence.
5
Plot learning curves first. If performance has plateaued, more data won't help.
6
Validate on the real distribution. If the test set doesn't match production, the results are fiction.

The approach

Most comparisons test one model on one dataset with one seed and declare a winner. I built a grid that varies architecture, pre-training, features, data scale, and random seed independently. Not to find the best model, but to find when each model is the right choice.

XGB
XGBoost
Tree-based ensemble
Gradient-boosted decision trees. Learns via feature splits on numerical values.
Click for details
FT-T
FT-Transformer
Tabular transformer
Transformer designed for tabular data. Per-feature tokenization, no positional encoding.
Click for details
BERT
BERT
Pre-trained language model
110M parameter transformer pre-trained on text. Requires serializing data to language.
Click for details

Why it's in the evaluation: XGBoost represents the established baseline for tabular data. Tree-based splits are optimized for numerical feature spaces and handle class imbalance natively via sample weighting. It's what most practitioners reach for first, and for good reason.

Key architectural property: Requires manual feature engineering. The model operates on whatever features it receives. It cannot discover semantic relationships or learn representations from raw data. Performance depends entirely on feature engineering quality.

Training speed: 0.4 to 5 seconds. Orders of magnitude faster than either transformer. At 360x faster than BERT, the speed advantage is itself a reason to prefer XGBoost when performance is close.

What the evaluation revealed: Dominated on PCA features (0.865). On semantic features, required three stages of manual engineering to reach 0.759, or 98% of BERT's performance with significant domain expertise investment.

Why it's in the evaluation: FT-Transformer was designed specifically for tabular data. It represents the hypothesis that architecture-data alignment matters more than model scale or pre-training. It's the purpose-built option.

Key architectural property: Each feature gets its own embedding (feature tokenization), and there's no positional encoding, because tabular features have no inherent order. This is a deliberate design choice: the architecture encodes the assumption that features are independent, which is true for tabular data and false for text.

What the evaluation revealed: Zero catastrophic failures across 75 experiments. 8x lower variance than BERT. Consistent mid-tier performance (0.689 to 0.729) without any feature engineering. Not the highest-performing model, but the most reliable. When architecture matches data structure by design, stability follows.

Why it's in the evaluation: BERT represents the "use the popular model" hypothesis. It dominates NLP benchmarks and is increasingly applied to tabular tasks by serializing features into text. The question is whether its pre-trained language understanding transfers to non-language data.

Key architectural property: 110M parameters pre-trained on Wikipedia and BookCorpus. Positional encoding assumes sequential token dependencies, an inductive bias that fits language but not independent tabular features. Requires converting structured data into text strings.

What the evaluation revealed: Won on genuine semantic features (0.803) where its pre-trained embeddings captured relationships like "Walmart" similar to "Target." Failed on numerical/PCA features (0.686) where its text-oriented architecture was a liability. A 4% catastrophic failure rate across seeds. High variance, high ceiling, high risk.

The evaluation grid
3
architectures
click
×
3
datasets
click
×
5
data scales
click
×
5
random seeds
click
→
197
total experiments
Three fundamentally different modeling approaches
XGBoost: tree-based, 0.4 to 5s training
FT-Transformer: tabular-native, ~3 min
BERT: pre-trained LM, ~30 min
Each with different feature types and scale
European: 284K, PCA numerical
Sparkov: 1M, semantic + categorical
Sparkov-matched: 36K, size control
Testing whether more data always helps
5%
10%
25%
50%
100%
Isolating model variance from lucky initialization
Seed 42
Seed 123
Seed 456
Seed 789
Seed 1011
Not every combination was tested (some model-dataset pairs were not applicable, some runs used 4 seeds). 197 is the actual count of completed experiments, each with controlled variation across these four dimensions.

The problem

Fraud detection is a highly imbalanced classification problem. Legitimate transactions outnumber fraudulent ones by several hundred to one, meaning a model can achieve over 99% accuracy by predicting the majority class for every input. The data exists in multiple representations (PCA-transformed numerical, raw categorical, semantic text), which allows the same architecture to be evaluated against fundamentally different feature structures. If a modeling assumption doesn't hold, this problem exposes it.

EU
European credit card
284,807 transactions · 2 days
576:1 class imbalance
30 PCA features Numerical Pre-engineered
Click for details
SK
Sparkov synthetic
1,296,675 transactions · 2 years
580:1 class imbalance
23 raw features Semantic text Categorical
Click for details

Why this dataset matters for the evaluation: The features are PCA-transformed, already pre-engineered into dense numerical representations. This removes the feature engineering variable entirely. If a model wins here, it's because of how it handles numerical feature splits, not because it's better at extracting signal from raw data.

Key characteristics: Anonymous features (V1 to V28) plus transaction amount and time. Only 492 fraudulent transactions in the entire dataset. Two days of data means limited temporal diversity, so the model can't learn seasonal fraud patterns.

What the evaluation revealed: XGBoost dominated (0.865) because tree-based splits are optimized for exactly this kind of pre-engineered numerical space. BERT's text-oriented architecture was a poor fit, scoring 0.686, nearly 18 points lower. Performance peaked at 25% of the data and flatlined, suggesting the PCA features saturate quickly.

Why this dataset matters for the evaluation: The features are raw and diverse: merchant names, transaction categories, cities, customer demographics. This is where semantic understanding should help, and where feature engineering effort becomes a variable. You can test BERT with text, FT-Transformer with categorical encodings, and XGBoost with hand-engineered features, all from the same underlying data, three different representations.

Key characteristics: 23 features spanning temporal, geographic, demographic, and transactional dimensions. Two years of data means real temporal diversity. 2,230 fraudulent transactions, enough rare patterns that downsampling risks losing them (which is exactly what happened with the size-matched control).

What the evaluation revealed: BERT won (0.803) by leveraging semantic relationships in merchant names and categories. But XGBoost with three stages of manual feature engineering reached 0.759 (98% of BERT's performance) at 360x the speed. FT-Transformer hit 0.689 with zero engineering. The tradeoff: automatic learning vs. domain expertise vs. speed.

Findings

Each card below pairs a standard ML practice with the precondition the evaluation grid revealed. The practice is real and well-supported. The precondition is what determines whether it holds for a given problem.

Best practice 1: Use deep learning
The practice
BERT outperforms traditional ML
Backed by benchmarks across NLP and increasingly across tabular tasks. Widely adopted as a strong default.
The precondition
Only when features are semantic
BERT won on semantic features (0.773). On PCA-transformed numerical features, XGBoost dominated (0.865). The feature type, not the model, determined the winner.
BERT/semantic: 0.773 · XGBoost/PCA: 0.865
Feature type determines the model hierarchy
Same models, different data. The winner flips completely.
EU
European credit card
284K transactions · 30 PCA-transformed numerical features · XGBoost wins
XGBoost0.865
FT-Transformer0.729
BERT0.686
SK
Sparkov synthetic
1.3M transactions · 23 raw features with merchant names, categories, geography · BERT wins
BERT0.803
XGBoost (3-stage engineering)0.759
FT-Transformer0.689
PCA numerical features favor tree-based splits. Semantic text features favor pre-trained language understanding. The data determines the winner.
XGBoost
FT-Transformer
BERT
Best practice 2: Use domain-specific pre-training
The practice
Domain pre-training improves transfer
Pre-training on domain-relevant data is a proven strategy. BioBERT, FinBERT, CodeBERT all show gains from domain alignment. The expectation: pre-training on fraud data should help with fraud detection.
The precondition
Only when the domains are structurally compatible
I ran a cross-domain transfer study: pre-trained on IEEE-CIS (US e-commerce, 3.5% fraud rate), fine-tuned on European (EU card-present, 0.17% fraud rate). Same task label, but 20x difference in fraud prevalence, entirely different feature spaces, different transaction types. The result was scale-dependent: benefit at small data, collapse at full data. A single-scale evaluation would have missed the whole story.
Peak at 10%: +14.7% over vanilla BERT · At 100%: 0.127 (collapse)
See the mechanism ↓

Pre-training source: IEEE-CIS fraud dataset (590K transactions, 3.5% fraud rate, 434 features, US online payments).

Fine-tuning target: European credit card data (284K transactions, 0.17% fraud rate, 30 PCA features).

Why it collapsed: The two datasets differ on every dimension. IEEE has a 3.5% fraud rate; European has 0.17% (20x rarer). IEEE has 434 raw categorical and transactional features; European has 30 PCA-transformed numerical features. IEEE covers US online payments; European covers EU card-present transactions. At small data scales, the IEEE representations still provided useful signal because the model had encountered fraud patterns before, even structurally different ones. At full data, there was enough gradient to overwrite those representations entirely, but the overwrite was unstable because the target fraud distribution, feature space, and transaction domain were all fundamentally different from what the representations were optimized for. The classification head learned to suppress fraud probability regardless of input.

Cross-domain transfer: the pipeline that produced Fraud-BERT
I took bert-base-uncased, pre-trained it on IEEE-CIS fraud data via masked language modeling, then fine-tuned on the European dataset. Both are fraud detection. The domains are structurally different.
Step 1: MLM pre-training
IEEE-CIS fraud dataset
Fraud rate 3.5%
Features 434
Transactions 590K
Domain US online
→
BERT
bert-base
110M params
→
Step 2: Fine-tuning
European credit card
Fraud rate 0.17%
Features 30
Transactions 284K
Domain EU cards
20×
fraud rate difference
14×
fewer features
different
geography + transaction type
Same task label ("fraud detection"), structurally different domains. At small data scales, pre-trained representations helped.
At full data, the fine-tuning gradient overwrote them unstably. Performance peaked at 10%, then collapsed.
Cross-domain transfer: benefit and harm are both real, depending on scale
PR-AUC by data scale, European dataset. Fraud-BERT peaks at 10% (where pre-trained representations help) then collapses at full data (where fine-tuning overwrites them).
XGBoost
BERT (generic)
Fraud-BERT (IEEE)
5%
10%
25%
50%
100%
XGBoost (steady improvement)
BERT (generic, no domain pre-training)
Fraud-BERT (IEEE fraud pre-training)
Note: two different failure modes. BERT's drop at 100% is partly driven by seed instability (one catastrophic seed pulls the mean down). Without the failed seed, BERT scores ~0.541. Fraud-BERT's collapse is a different mechanism entirely: 3 of 4 seeds at 100% produced classifiers with zero discriminative power. The cross-domain pre-training became destructive once fine-tuning data had enough gradient to overwrite it.
The general lesson: A shared task label ("fraud detection") is not a reliable proxy for transfer compatibility. The portable insight applies anywhere someone is considering domain pre-training: check structural alignment (prevalence, feature space, domain characteristics), not just task name.
Four more preconditions the grid revealed
Best practice 3
Pre-trained embeddings extract rich features
Precondition: Only when features are genuinely textual. Same model on real merchant names: 0.803. On numeric features serialized to text: 0.706.
0.803
genuine text
0.706
numeric to text
Best practice 4
More data improves performance
Precondition: Only when the data adds new information. European peaked at 25% and flatlined. 75% more data, 0% improvement. Sparkov improved to 100%.
0.746
European at 25%
0.729
European at 100%
Best practice 5
A good validation score means the model works
Precondition: Only if stable across initializations. BERT seed 456: total class collapse, caught 3 of 492 fraud cases. 4% catastrophic failure rate, invisible to a single run.
0.509 to 0.651
4 normal seeds
0.032
seed 456 (collapse)
Best practice 6
Size-match datasets for fair comparison
Precondition: Only if downsampling preserves diversity. Stratified sampling kept the fraud rate but reduced 1,720 fraud patterns to 62. The methodology produced false confidence.
0.721
validation
0.346
test (50% drop)
One seed in 24: complete class collapse
BERT PR-AUC across random seeds, European dataset, 100% data
Seed 42
0.564
Seed 123
0.509
Seed 456
0.032
Seed 789
0.651
Seed 456 predicted 99.8% of transactions as legitimate, catching 3 of 492 fraud cases. The model ran to completion with decreasing training loss, giving no indication of failure during training. In practice, a result this bad wouldn't be shipped. But if this were the only seed run, it would have muddied the entire evaluation: is the architecture wrong, the features wrong, the hyperparameters wrong, or was it just an unlucky initialization? Running multiple seeds from the start eliminates that confusion entirely.
Synthesis

XGBoost vs FT-Transformer: the real tradeoff

On structured tabular data (numerical and categorical features), the practical comparison is between XGBoost and FT-Transformer. BERT is not competitive in this context. The question is whether the performance ceiling or the reliability floor matters more for a given deployment.

XGB
XGBoost
Highest ceiling on both datasets. Requires manual feature engineering on raw features.
European (PCA provided) 0.865
Sparkov (manual engineering) 0.759
Training time 0.4 to 5s
Failure rate 0%
Feature engineering 3 stages required
Sparkov engineering progression:
Stage 1 (numerical only): 0.109
Stage 2 (+categorical encoding): 0.559
Stage 3 (+behavioral/geospatial): 0.759
Each stage required domain knowledge to identify which features to build.
FT-T
FT-Transformer
Consistent performance on raw features with no manual engineering. Architecture learns feature relationships automatically.
European (PCA provided) 0.729
Sparkov (raw categorical) 0.689
Training time ~3 min
Failure rate 0%
Feature engineering None beyond raw input
The tradeoff: Lower ceiling (0.729 vs 0.865 on European, 0.689 vs 0.759 on Sparkov) but zero engineering effort and the lowest variance of any model tested. 8x more stable than BERT across seeds.
EU
On structured tabular data
XGBoost operates on whatever features it receives; it cannot discover relationships independently. FT-Transformer learns feature importance via per-feature embeddings. The tradeoff is domain expertise vs. automatic learning at a lower ceiling. XGBoost wins on performance when engineered well. FT-Transformer wins on reliability with zero effort.
SK
On semantic text features
BERT dominated at 0.803 by leveraging pre-trained embeddings that capture relationships between merchant names, categories, and transaction descriptions. XGBoost with manual engineering reached 0.759 (98% of BERT, 360x faster). FT-Transformer scored 0.689 with no engineering. When the data is genuinely textual, pre-trained language understanding provides a ceiling that tabular-native architectures cannot match.

The workflow this implies

Start with the problem, not the model:

1
Start with data characteristics, not model selection. Class distribution, feature types, data volume, diversity of the target class. The data tells you which practices are likely to apply.
2
Match architecture to feature representation. BERT's strength is semantic understanding. XGBoost's strength is numerical feature splits. The architecture should fit the data, not the other way around.
3
Test cross-domain transfer at multiple scales. A shared task label does not guarantee structural compatibility. Run the comparison, and check whether the result changes with data volume.
4
Run multiple seeds from the start. 3 to 5 seeds catches rare failures and gives you confidence intervals. A single score is an anecdote, not evidence.
5
Plot learning curves before scaling data collection. If performance has plateaued, more data won't help. Invest in better features or architecture instead.
6
Validate on the real distribution. If your test set doesn't match production conditions, your results describe a world that doesn't exist.

Every best practice here works, under conditions determined by the problem. The skill is checking whether those conditions are met before committing. That's what this evaluation system does, and it's the same skill that matters for any AI system in production.