Topic: training data
4 stories found
Yesterday
Adaption Labs Introduces βInvent a Datasetβ: Training Data Generated From a Task Description, Not a Seed Corpus
Adaption Labs has launched Invent a Dataset, a tool that creates training data based on task descriptions without needing a seed corpus or manual labeling, aiming to streamline the training process for machine learning models. This innovation could significantly reduce the time and effort required in preparing datasets, making model development more accessible and efficient.
Friday, September 4, 2026
Contamination Inflates Scores but Rarely Reorders Large Language Model Leaderboards
Benchmark contamination can inflate scores by leaking test items into training data, but its impact on reordering LLM leaderboards is limited, suggesting the reliability threat may be overstated.
Tuesday, August 25, 2026
Mitigating Bias in Large Vision-Language Models via Counterfactual Ensemble Decoding
Large Vision-Language Models (LVLMs), while effective in various tasks, can exhibit biased behavior due to social biases in their training data. Researchers propose a method called Counterfactual Ensemble Decoding to mitigate these biases, highlighting the importance of addressing fairness in AI systems.
πΏ That's all for now. Come back tomorrow.
4 of 4 items shown. Sources: 107 days indexed.