Skip to content

Association discovery

Association discovery encompasses a broad family of related tasks that identify key interactions between entities or elements within data. These include association mining, pattern mining, association rule discovery, subgroup discovery, emerging pattern discovery and contrast discovery.

My pioneering association discovery techniques seek the most useful associations, rather than applying the minimum-support constraint more commonly used in the field.  Many of these techniques are included in my Magnum Opus software, which is now incorporated in BigML.  Magnum Opus has been widely used in scientific research.

Due to the large numbers of patterns that are considered in association discovery, most techniques suffer large risk of type-1 error. That is, they are likely to find patterns that appear interesting only due to chance artifacts of the process by which the sample data were generated.  Most attempts to control this risk do so at the cost of high risk of type-2 error. That is, they are likely to falsely reject non-spurious patterns.  I have pioneered strategies for statistically sound pattern discovery, strictly controlling type-1 error during association discovery without the level of risk of type-2 error suffered by previous approaches.

The OPUSMiner statistically sound itemset discovery software can be downloaded here. An R package is available here.

The Skopus statistically sound sequential pattern discovery software can be downloaded here.

An implementation of impact rules can be downloaded here: https://github.com/501856869/Significant_impact_rule.

ACM SIGKDD 2014 Tutorial on Statistically Sound Pattern Discovery, with Wilhelmiina Hämäläinen.

Publications

Boley, Mario; Teshuva, Simon; Bodic, Pierre Le; Webb, Geoffrey I

Better Short than Greedy: Interpretable Models through Optimal Rule Boosting

Proceedings of the 2021 SIAM International Conference on Data Mining (SDM), pp. 351-359, SIAM 2021.

Abstract | Links | BibTeX

Hamalainen, Wilhelmiina; Webb, Geoffrey I.

A tutorial on statistically sound pattern discovery

Data Mining and Knowledge Discovery, vol. 33, no. 2, pp. 325-377, 2019, ISSN: 1573-756X.

Abstract | Links | BibTeX

Shi, Wenzhong; Zhang, Anshu; Webb, Geoffrey I.

Mining significant crisp-fuzzy spatial association rules

International Journal of Geographical Information Science, vol. 32, no. 6, pp. 1247-1270, 2018.

Abstract | Links | BibTeX

Hamalainen, Wilhelmiina; Webb, Geoffrey I

Specious rules: an efficient and effective unifying method for removing misleading and uninformative patterns in association rule mining

Proceedings of the 2017 SIAM International Conference on Data Mining, pp. 309-317, SIAM 2017.

BibTeX

Petitjean, Francois; Li, Tao; Tatti, Nikolaj; Webb, Geoffrey I.

Skopus: Mining top-k sequential patterns under leverage

Data Mining and Knowledge Discovery, vol. 30, no. 5, pp. 1086-1111, 2016, ISSN: 1573-756X.

Abstract | Links | BibTeX

Webb, Geoffrey I.; Petitjean, Francois

A multiple test correction for streams and cascades of statistical hypothesis tests

Proceedings of the ACM SIGKDD Conference on Knowledge Discovery and Data Mining, KDD-16, pp. 1255-1264, ACM Press, 2016.

Abstract | Links | BibTeX

Zhang, Anshu; Shi, Wenzhong; Webb, Geoffrey I.

Mining significant association rules from uncertain data

Data Mining and Knowledge Discovery, vol. 30, no. 4, pp. 928-963, 2016.

Abstract | Links | BibTeX

Petitjean, F.; Webb, G. I.

Scaling log-linear analysis to datasets with thousands of variables

Proceedings of the 2015 SIAM International Conference on Data Mining, pp. 469-477, 2015.

Abstract | Links | BibTeX

Petitjean, F.; Allison, L.; Webb, G. I.

A Statistically Efficient and Scalable Method for Log-Linear Analysis of High-Dimensional Data

Proceedings of the 14th IEEE International Conference on Data Mining, pp. 480-489, 2014.

Abstract | Links | BibTeX

Webb, G. I.; Vreeken, J.

Efficient Discovery of the Most Interesting Associations

ACM Transactions on Knowledge Discovery from Data, vol. 8, no. 3, 2014.

Abstract | Links | BibTeX

Hamalainen, Wilhelmiina; Webb, Geoffrey I.

Statistically sound pattern discovery

KDD '14: Proceedings of the 20th ACM SIGKDD international conference on Knowledge discovery and data mining, pp. 1976, 2014.

Abstract | Links | BibTeX

Petitjean, F.; Webb, G. I.; Nicholson, A. E.

Scaling log-linear analysis to high-dimensional data

Proceedings of the 13th IEEE International Conference on Data Mining, pp. 597-606, 2013.

Abstract | Links | BibTeX

Webb, G. I.

Filtered-top-k Association Discovery

WIREs Data Mining and Knowledge Discovery, vol. 1, no. 3, pp. 183-192, 2011.

Abstract | Links | BibTeX

Webb, G. I.

Self-Sufficient Itemsets: An Approach to Screening Potentially Interesting Associations Between Items

ACM Transactions on Knowledge Discovery from Data, vol. 4, iss. 1, 2010.

Abstract | Links | BibTeX

Novak, P.; Lavrac, N.; Webb, G. I.

Supervised Descriptive Rule Discovery: A Unifying Survey of Contrast Set, Emerging Pattern and Subgroup Mining

Journal of Machine Learning Research, vol. 10, pp. 377-403, 2009.

Abstract | Links | BibTeX

Webb, G. I.

Layered Critical Values: A Powerful Direct-Adjustment Approach to Discovering Significant Patterns

Machine Learning, vol. 71, no. 2-3, pp. 307-323, 2008.

Abstract | Links | BibTeX

Webb, G. I.

Discovering Significant Patterns

Machine Learning, vol. 68, no. 1, pp. 1-33, 2007.

Abstract | Links | BibTeX

Butler, S.; Webb, G. I.

Mining Group Differences

Wang, John (Ed.): The Encyclopedia of Data Warehousing and Mining, pp. 795-799, Idea Group Inc., Hershey, PA, 2006.

Links | BibTeX

Huang, S.; Webb, G. I.

Efficiently Identifying Exploratory Rules' Significance

LNAI State-of-the-Art Survey series, 'Data Mining: Theory, Methodology, Techniques, and Applications', pp. 64-77, Springer, Berlin/Heidelberg, 2006, (An earlier version of this paper was published in S.J. Simoff and G.J. Williams (Eds.), Proceedings of the Third Australasian Data Mining Conference (AusDM04) Cairns, Australia. Sydney: University of Technology, pages 169-182.).

Abstract | Links | BibTeX

Webb, G. I.

Discovering Significant Rules

Ungar, L.; Craven, M.; Gunopulos, D.; Eliassi-Rad, T. (Ed.): Proceedings of the Twelfth ACM SIGKDD International Conference on Knowledge Discovery and Data Mining (KDD-2006), pp. 434-443, The Association for Computing Machinery, Philadelphia, PA, 2006.

Abstract | Links | BibTeX

Huang, S.; Webb, G. I.

Pruning Derivative Partial Rules During Impact Rule Discovery

Ho, T. B.; Cheung, D.; Liu, H. (Ed.): Lecture Notes in Computer Science Vol. 3518: Proceedings of the 9th Pacific-Asia Conference on Advances in Knowledge Discovery and Data Mining (PAKDD 2005), pp. 71-80, Springer, Hanoi, Vietnam, 2005.

Abstract | BibTeX

Huang, S.; Webb, G. I.

Discarding Insignificant Rules During Impact Rule Discovery in Large, Dense Databases

Kargupta, H.; Kamath, C.; Srivastava, J.; Goodman, A. (Ed.): Proceedings of the Fifth SIAM International Conference on Data Mining (SDM'05) [short paper], pp. 541-545, Society for Industrial and Applied Mathematics, Newport Beach, CA, 2005.

Abstract | BibTeX

Webb, G. I.

K-Optimal Pattern Discovery: An Efficient and Effective Approach to Exploratory Data Mining

Zhang, S.; Jarvis, R. (Ed.): Lecture Notes in Computer Science 3809: Advances in Artificial Intelligence, Proceedings of the 18th Australian Joint Conference on Artificial Intelligence (AI 2005)[Extended Abstract], pp. 1-2, Springer, Sydney, Australia, 2005.

Links | BibTeX

Webb, G. I.; Zhang, S.

k-Optimal-Rule-Discovery

Data Mining and Knowledge Discovery, vol. 10, no. 1, pp. 39-79, 2005.

Abstract | Links | BibTeX

Thiruvady, D. R.; Webb, G. I.

Mining Negative Rules using GRD

Dai, H.; Srikant, R.; Zhang, C. (Ed.): Lecture Notes in Computer Science Vol. 3056: Proceedings of the Eighth Pacific-Asia Conference on Knowledge Discovery and Data Mining (PAKDD 04) [Short Paper], pp. 161-165, Springer, Sydney, Australia, 2004.

Abstract | BibTeX

Webb, G. I.

Preliminary Investigations into Statistically Valid Exploratory Rule Discovery

Simoff, S. J.; Williams, G. J.; Hegland, M. (Ed.): Proceedings of the Second Australasian Data Mining Conference (AusDM03), pp. 1-9, University of Technology, Canberra, Australia, 2003.

Abstract | BibTeX

Webb, G. I.

Association Rules

Ye, Nong (Ed.): The Handbook of Data Mining, Chapter 2, pp. 25-39, Lawrence Erlbaum Associates, 2003.

BibTeX

Webb, G. I.; Butler, S.; Newlands, D.

On Detecting Differences Between Groups

Domingos, P.; Faloutsos, C.; Senator, T.; Kargupta, H.; Getoor, L. (Ed.): Proceedings of The Ninth ACM SIGKDD International Conference on Knowledge Discovery and Data Mining (KDD-2003), pp. 256-265, The Association for Computing Machinery, Washington, DC, 2003.

Abstract | BibTeX

Zhang, C.; Zhang, S.; Webb, G. I.

Identifying Approximate Itemsets of Interest In Large Databases

Applied Intelligence, vol. 18, pp. 91-104, 2003.

Abstract | Links | BibTeX

Webb, G. I.; Zhang, S.

Removing Trivial Associations in Association Rule Discovery

Proceedings of the First International NAISO Congress on Autonomous Intelligent Systems (ICAIS 2002), NAISO Academic Press, Geelong, Australia, 2002.

Abstract | BibTeX

Webb, G. I.; Zhang, S.

Further Pruning for Efficient Association Rule Discovery

Stumptner, M.; Corbett, D.; Brooks, M. J. (Ed.): Lecture Notes in Computer Science Vol. 2256: Proceedings of the 14th Australian Joint Conference on Artificial Intelligence (AI'01), pp. 605-618, Springer, Adelaide, Australia, 2001.

Abstract | BibTeX

Webb, G. I.

Discovering Associations with Numeric Variables

Provost, F.; Srikant, R. (Ed.): Proceedings of the Seventh ACM SIGKDD International Conference on Knowledge Discovery and Data Mining (KDD-2001)[short paper], pp. 383-388, The Association for Computing Machinery, San Francisco, CA, 2001.

Abstract | Links | BibTeX

Webb, G. I.

Efficient Search for Association Rules

Ramakrishnan, R.; Stolfo, S. (Ed.): Proceedings of the Sixth ACM SIGKDD International Conference on Knowledge Discovery and Data Mining (KDD-2000), pp. 99-107, The Association for Computing Machinery, Boston, MA, 2000.

Abstract | BibTeX

Webb, G. I.

Inclusive Pruning: A New Class of Pruning Rule for Unordered Search and its Application to Classification Learning

Ramamohanarao, K. (Ed.): Australian Computer Science Communications Vol. 18 (1): Proceedings of the Nineteenth Australasian Computer Science Conference (ACSC'96), pp. 1-10, ACS, Royal Melbourne Insitute of Technology, Australia, 1996.

Abstract | BibTeX

Webb, G. I.

OPUS: An Efficient Admissible Algorithm For Unordered Search

Journal of Artificial Intelligence Research, vol. 3, pp. 431-465, 1995.

Abstract | Links | BibTeX