Skip to content

Bias and variance

Bias and variance provide a powerful conceptual tool for analyzing machine learning performance.  My research revealed that as sample size increases, the ratio of variance relative to bias tends to decrease. As has since become widely accepted, this implies that low variance learning algorithms, such as linear models, should be most effective with small data quantity. For large data quantities, low bias learning algorithms, such as deep learning, should be most effective.

Previous approaches to conducting bias-variance experiments have provided little control over the types of data distribution from which bias and variance are estimated.   I have developed new techniques for bias-variance analysis that provide greater control over the data distribution.  Experiments show that the type of distribution used for bias-variance experiments can greatly affect the results obtained.

Publications

Webb, Geoffrey I; Conilione, Paul

Estimating bias and variance from data

2004, (Unpublished manuscript).

Abstract | BibTeX

Brain, D.; Webb, G. I.

The Need for Low Bias Algorithms in Classification Learning From Large Data Sets

Lecture Notes in Computer Science 2431: Principles of Data Mining and Knowledge Discovery: Proceedings of the Sixth European Conference (PKDD 2002), pp. 62-73, Springer-Verlag, Helsinki, Finland, 2002.

Abstract | BibTeX

Webb, G. I.

MultiBoosting: A Technique for Combining Boosting and Wagging

Machine Learning, vol. 40, no. 2, pp. 159-196, 2000.

Abstract | Links | BibTeX

Brain, D.; Webb, G. I.

On The Effect of Data Set Size on Bias And Variance in Classification Learning

Richards, D.; Beydoun, G.; Hoffmann, A.; Compton, P. (Ed.): Proceedings of the Fourth Australian Knowledge Acquisition Workshop (AKAW-99), pp. 117-128, The University of New South Wales, Sydney, Australia, 1999.

Abstract | BibTeX