Seminar - [SMS seminar] On statistical analysis of grouped data
School of Mathematics and Statistics Research Seminar
Speaker: Estate Khmaladze, Professor Emeritus
Time:
Friday 7th August 2026 at 02:00 PM -
03:00 PM
Location:
Cotton Club,
Cotton 350
Groups:
"Mathematics"
"Statistics and Operations Research"
Abstract
[Due to the unforeseen circumstances, the seminar had to be moved. Thanks for your patience.]
The talk represents part of the material from the speaker's joint work with Sara Algeri (U.Minnesota). The talk could have been called ``Statistical analysis of large number of infrequent events", or ``On theory of Poisson regression", or ``When Pearson's chi^2 is not a goodness of fit test". Whatever it is called, it is clear we will start with well trodden topic of statistics - and if we think, as the speaker did, that we have enough knowledge about the phenomena there, or enough intuition about the subject, we may be not entirely correct.
The talk is our attempt to say something useful for everyday statistical work. The key words are:
divisible statistics, fine partitioning, large number of groups (or urns, boxes, pixels, etc.), contiguous alternatives, useless parts of useful tests, real astronomical data, testing uniformity, spectral statistics, statistics of empty boxes.
Sara Algeri, Estate V. Khmaladze (2026), On the statistical analysis of grouped data: when Pearson chi^2 and other divisible statistics are not goodness-of-fit, JRSS B, September 26, online - 23 June 2026.
===An updated/extended abstract:======
The previous abstract is valid - but 1 hour is not enough.
So, I would much rather say and illustrate one thing, but clearly - good enough for p/g students, than to cover large part of S.Algeri-Khm in more telegraphic manner.
Therefore, out of five topics below let me, please, describe and illustrate mostly B. -- but with better clarity.
How to create a unified approach to the theory: not \chi^2 separately and spectral statistics separately, as "similar but different", but as the same object - linear functional from a single empirical process. Why for any given divisible statistic we can construct another divisible statistic with uniformly better power against all contiguous alternatives. For example, why one can construct a test, uniformly better than K.Pearson's \chi^2? When the hypothesis is parametric, one will need to estimate the parameter. However, frequently expressed "We'd like to know exact value, but unfortunately..." is incorrect attitude, because estimation of parameter gives an advantage in power. When we need to estimate parameter, based on divisible statistic, do we know what is the best estimating equation? - We do, which is a Cram\'er-Rao type result. One cannot have a goodness of fit(GoF) test based on one or "few" divisible statistics. But can we have a GoF test based on partial sums' process? -- Yes, we can construct a class of such processes, and they will lead to GoF tests with analytically known limit distributions.