Studying people and computers (www.nickmvincent.com) Blogging about data and steering AI (dataleverage.substack.com)
dc: data minimization
Part of the datacounterfactuals.org reading lists. Research on data minimization: reducing data collection, retention, and feature disclosure while preserving utility, with links across privacy by design, GDPR, ML, experimentation, scaling laws, and data streams.
19 cards
2mo
upcoming conferences
I try to record conferences that I, or my students, might be interested in. I'll try to move the "done" conferences to a different collection at some regular interval.
3 cards
2mo
completed conferences of interest
1 card
2mo
dc: data provenance and source attribution
Part of the datacounterfactuals.org reading lists. Research on data provenance, dataset documentation, licensing and attribution audits, and technical source-attribution methods for understanding which data sources are available, permitted, or responsible for model behavior.
5 cards
3mo
dc: user-generated content
Part of the datacounterfactuals.org reading lists. Research on the value of user-generated content for AI systems, search engines, and digital platforms. Explores how content from Wikipedia, social media, and other user contributions powers modern AI and what this means for content creators.
9 cards
3mo
dc: scaling laws
Part of the datacounterfactuals.org reading lists. Research on scaling laws describing how model performance changes with data, parameters, and compute. Foundational work for understanding data requirements and compute-optimal training.
6 cards
3mo
dc: fairness via data interventions
Part of the datacounterfactuals.org reading lists. Research on fairness through data-centric approaches.
7 cards
3mo
dc: active learning
Part of the datacounterfactuals.org reading lists. Research on active learning: selecting which data points to label next given a limited annotation budget. Connects to coresets, curriculum learning, and efficient data collection strategies.
6 cards
3mo
dc: causality
Part of the datacounterfactuals.org reading lists. Foundational research on causal inference and experimental design.
5 cards
3mo
dc: influence
Part of the datacounterfactuals.org reading lists. Research on influence functions and data attribution methods. Covers techniques for understanding which training examples are responsible for model predictions and behaviors.
19 cards
3mo
dc: semivalues
Part of the datacounterfactuals.org reading lists. Research on data valuation methods, including Data Shapley, semivalue variants like Beta Shapley and Data Banzhaf, and related approaches for quantifying individual training point contributions.
27 cards
3mo
dc: training dynamics
Part of the datacounterfactuals.org reading lists. Research on using training dynamics to diagnose, rank, or explain examples.
4 cards
3mo
dc: selection and coresets
Part of the datacounterfactuals.org reading lists. Research on coreset selection and efficient data subset methods.
4 cards
3mo
dc: unlearning
Part of the datacounterfactuals.org reading lists. Research on machine unlearning: updating models as if specific training data had never been included.
7 cards
4mo
dc: poisoning
Part of the datacounterfactuals.org reading lists. Research on data poisoning attacks and defenses.
5 cards
4mo
www.aies-conference.com
AI, Ethics, and Society — Home
sigchi.org
CI 2026
neurips.cc
Call for Papers 2026
aclanthology.org
WASA: WAtermark-based Source Attribution for Large Language Model-Generated Data
proceedings.iclr.cc
SILO Language Models: Isolating Legal Risk In a Nonparametric Datastore
openreview.net
LAION-5B: An open large-scale dataset for training next generation image-text models
www.jmlr.org
Exploring the Limits of Transfer Learning with a Unified Text-to-Text Transformer
aclanthology.org
Documenting Large Webtext Corpora: A Case Study on the Colossal Clean Crawled Corpus
aclanthology.org
DBpedia Abstracts: A Large-Scale, Open, Multilingual NLP Training Corpus
proceedings.neurips.cc
www.cs.cmu.edu
Active Learning with Statistical Models

link.springer.com
A Sequential Algorithm for Training Text Classifiers
openreview.net
Deep Batch Active Learning by Diverse, Uncertain Gradient Lower Bounds
doi.org
Query by committee
burrsettles.com
iopscience.iop.org
Deep double descent: where bigger models and more data hurt

arxiv.org
Beyond neural scaling laws: beating power law scaling via data pruning

arxiv.org
Scaling Laws for Autoregressive Generative Modeling

arxiv.org
Deep Learning Scaling is Predictable, Empirically

arxiv.org
Scaling Laws for Neural Language Models

arxiv.org
Training Compute-Optimal Large Language Models

arxiv.org
From Principle to Practice: Vertical Data Minimization for Machine Learning

arxiv.org
Configurable Per-Query Data Minimization for Privacy-Compliant Web APIs

arxiv.org
Operationalizing the Legal Principle of Data Minimization for...

arxiv.org
Monitoring Data Minimisation

arxiv.org
Data Minimisation: a Language-Based Approach (Long Version)

arxiv.org
Data Minimisation in Communication Protocols: A Formal Analysis...
proceedings.mlr.press
Characterizing Fairness Over the Set of Good Models Under Selective Labels

doi.org
Certifying and Removing Disparate Impact
proceedings.neurips.cc
Optimized Pre-Processing for Discrimination Prevention
proceedings.mlr.press
Learning Fair Representations
link.springer.com
Data preprocessing techniques for classification without discrimination

arxiv.org
Datasheets for Datasets
aclanthology.org
WASA: WAtermark-based Source Attribution for Large Language Model-Generated Data
proceedings.iclr.cc
SILO Language Models: Isolating Legal Risk In a Nonparametric Datastore

arxiv.org
Datasheets for Datasets
aclanthology.org
Documenting Large Webtext Corpora: A Case Study on the Colossal Clean Crawled Corpus

www.nature.com
A large-scale audit of dataset licensing and attribution in AI
doi.org
Eternal Sunshine of the Spotless Net: Selective Forgetting in Deep Networks
aclanthology.org
Unlearning Traces the Influential Training Data of Language Models
papers.nips.cc
Making AI Forget You: Data Deletion in Machine Learning
ieeexplore.ieee.org
Towards Making Systems Forget with Machine Unlearning
proceedings.mlr.press
Descent-to-Delete: Gradient-Based Methods for Machine Unlearning
proceedings.mlr.press
Certified Data Removal from Machine Learning Models
openreview.net
Data Valuation Without Training of a Model
proceedings.neurips.cc
Deep Learning on a Data Diet: Finding Important Examples Early in Training
www.microsoft.com
An Empirical Study of Example Forgetting during Deep Neural Network Learning
aclanthology.org
Dataset Cartography: Mapping and Diagnosing Datasets with Training Dynamics

arxiv.org
Bayesian Influence Functions for Hessian-Free Data Attribution
openreview.net
Rescaled Influence Functions: Accurate Data Attribution in High...
openreview.net
Taming Hyperparameter Sensitivity in Data Attribution: Practical...
openreview.net
Distributional Training Data Attribution: What do Influence...
proceedings.neurips.cc
Stochastic Amortization: A Unified Approach to Accelerate Feature and Data Attribution
proceedings.neurips.cc
Training Data Attribution via Approximate Unrolling
proceedings.mlr.press
Collaborative Causal Inference with Fair Incentives
dash.harvard.edu
Estimating Causal Effects of Treatments in Randomized and Nonrandomized Studies
dash.harvard.edu
Reducing Bias in Observational Studies Using Subclassification on the Propensity Score

academic.oup.com
The central role of the propensity score in observational studies for causal effects

doi.org
Causality
conference2026.eaamo.org
EAAMO Conference 2026
proceedings.mlr.press
Rethinking Data Shapley for Data Selection Tasks: Misleads and Merits
openreview.net
Active Learning for Convolutional Neural Networks: A Core-Set Approach
cdn.aaai.org
proceedings.mlr.press
Coresets for Data-efficient Training of Machine Learning Models

arxiv.org
Semivalue-based data valuation is arbitrary and gameable
openreview.net
On the Impact of the Utility in Semivalue-based Data Valuation
proceedings.iclr.cc
Data Shapley in One Training Run
proceedings.neurips.cc
Robust Data Valuation with Weighted Banzhaf Values
openreview.net
SAVA: Scalable Learning-Agnostic Data Valuation
proceedings.mlr.press
An Instrumental Value for Data Production and its Application to Data Pricing

arxiv.org
Targeted Backdoor Attacks on Deep Learning Systems Using Data Poisoning
openreview.net
Backdoor or Feature? A New Perspective on Data Poisoning
proceedings.neurips.cc
Certified Defenses for Data Poisoning Attacks

arxiv.org
BadNets: Identifying Vulnerabilities in the Machine Learning Model...

arxiv.org
Poisoning Attacks against Support Vector Machines