Group Cross-Correlations with Faintly Constrained Filters
Abstract
Group convolutional layers with respect to some group are modeled by convolutions or cross-correlations with a filter, and they provide the fundamental building block for group convolutional neural networks. For entirely unconstrained filters and a non-abelian group, any hidden layer of such a network requires as many nodes as vertices in a fine enough discretization of . In order to reduce the necessary number of nodes, certain constraints on filters were proposed in the literature. We propose weaker constraints retaining this benefit while also resolving an incompatibility previous constraints have for group actions with non-compact stabilizers. Moreover, we generalize previous results to group actions that are not necessarily transitive, and we weaken the common assumption that is unimodular.
1 Introduction
Group convolutional neural networks with respect to some group were introduced by Cohen and Welling (2016). The distinguishing feature of such a neural network is that some of its layers are modeled by cross-correlations of functions with some filter henceforth denoted by (or ). Initially, Cohen and Welling (2016) impose no constraints on the filter (there denoted by ) and as a result, within any hidden layer feature vectors are modeled by functions defined on the entire group rather than a quotient of when is non-abelian. Kondor and Trivedi (2018, Definition 5) resolved this limitation by imposing a constraint on (there denoted by ), which can be paraphrased as bi-invariance. Cohen et al. (2019) generalized cross-correlations from transformations of vector-valued functions to transformations of sections to -equivariant vector bundles. In order to accommodate the -action on total spaces, the bi-invariance constraint of (Kondor and Trivedi, 2018, Definition 5) is generalized to their constraint of bi-equivariance. As we will see in Section 4.1.2, this constraint of bi-equivariance (or bi-invariance) is too strict, when stabilizers are not compact. In the present paper we propose a weaker constraint on with (24) that also works for non-compact stabilizers. Thinking of bi-equivariance as having both left-equivariance and right-equivariance, our constraint can be paraphrased as “equivariance with respect to conjugation”, which is implied by bi-equivariance.
Other than the compactness of stabilizers, a common assumption is that the considered -action should be transitive. In our Definition 2.4 of a cross-correlation, transitivity is not imposed. We also weaken the assumption that is unimodular. Otherwise, we adopt the setting and the generality by Cohen et al. (2019). In order to accommodate non-transitive -actions when comparing cross-correlations to integral transforms, we introduce orbitwise integral transforms in Section 3 and we recover the widely known correspondence between -equivariant integral transforms and -equivariant kernels henceforth denoted by . Finally, we establish a close relation between -equivariant orbitwise integral transforms and cross-correlations in Section 4.
2 Group Cross-Correlations
We start with a concrete example which will then be generalized so as to obtain a rather broad notion of a cross-correlation with Definition 2.4. As a group we consider the real numbers with addition and as an -space we consider the unit circle endowed with the -action
| (1) |
Now suppose we meant to design a neural network layer having continuous functions as inputs and as outputs. As inputs such functions could for example describe temperatures measured in each point or velocities in counter-clockwise direction of a fluid constrained to the circle. Moreover, suppose that the receptive field of each point should be limited to a small neighborhood. Such a transformation of a function can be obtained as a cross-correlation with a filter supported on a small neighborhood of :
| (2) |
Here we added the hat to the -operator as we will use the unaltered -operator to denote the general form of cross-correlations provided by Definition 2.4.
The action (1) also induces an -action on functions
where is the rotation of (as its graph) by in counter-clockwise direction, i.e. we have
| (3) |
for all , , and . Note that in the present paper function application takes precedence over the binary operator “” denoting group actions. Moreover, the operation
is -equivariant, i.e. for any , continuous , and we have the equation
| (4) |
or equivalently
| (5) |
2.1 Generalization to Non-Abelian Groups
Now let be a Hausdorff topological group with its neutral element and let be a Hausdorff -space throughout, i.e. we have a continuous -action
(We use the letter “” as it will also serve as a base space for vector bundles from Section 2.2 onward.) We now consider the not necessarily abelian group acting on in place of the additive group acting on in order to see where we get stuck, when we try to obtain a counterpart to equation (5) or equivalently -equivariance.
In analogy to the above example we assume we have a continuous compactly supported function and we also adapt the -operator from (2) by setting
| (6) |
for all continuous and , where is some locally finite Borel measure on . We now examine where the operation
fails to be -equivariant. The -action on functions is defined completely analogously as
where
Now let , , and . Then a counterpart to the left-hand side of (5) simplifies as
| (7) |
If was abelian (or if was in the center of ), then we could use to continue with a simplification similar to (4) ending in (6). In order to get rid of the conjugation by in the last term of (7), we make both the measure for integration and the filter dependent on the argument provided to cross-correlations, here in (6) and in (7).
So for the remainder of this paper excluding appendices we assume we have a family
of locally finite Borel measures that is compatible with the group action in the sense that
| (8) |
for any and , where is the pushforward measure of along the conjugation . For now, we also assume we have a continuous function
so as to provide a filter for any subject to the constraint
| (9) |
for all and . Moreover, in order to obtain well-defined cross-correlations, we assume each of the partially applied (filter) functions for to be compactly supported. Then for a continuous function , we define a cross-correlation by
| (10) |
Indeed, as a counterpart to (5) we obtain the equation
| (11) |
for all and . Here the third equality follows from the change of variables formula for the pushforward measure .
2.1.1 Continuity of Cross-Correlations
In order for the function defined in (10) to be continuous, we also assume the family to be continuous in the sense that
is a continuous function for any compactly supported continuous . Here the support of – denoted by – refers to the closure of the set and by saying that is compactly supported we mean that its support is compact. In the Appendix A.1 and specifically Theorem A.6 we show how such a family of measures can be constructed in a natural way.
2.1.2 Constraints as Stabilizer Invariances
Even though the filter now depends on a point as a second argument, the constraint (9) entails that the partially applied function is fully determined by for any and . In particular, if acts transitively on , then providing a filter as is the same as to provide just one of the partially applied functions for some satisfying the constraint
| (12) |
for all and . Here denotes the stabilizer of at .
If the group is abelian, then each partially applied function for is entirely unconstrained and we have for any two in the same orbit of the action . In this case, it comes down to a design choice, whether we use a filter as that is effectively parametrized by the orbits of (which are points of the quotient space ) or if we reuse the same weights provided by a single unconstrained filter across all orbits with the cross-correlation specified by (6).
In particular, if is abelian and it acts transitively on , then a cross-correlation as specified by (10) reduces to a cross-correlation as in (6) with a single unconstrained filter . This is in stark contrast with the previous approach to group convolutions, where there is a non-trivial constraint on filters as soon as there are non-trivial stabilizers, see for example (Kondor and Trivedi, 2018, Definition 5), (Cohen et al., 2019, Theorem 3.2), or (Gerken et al., 2023, Section 3.2).
In general, we may choose some fundamental domain so the filter is fully described by the family
of partially applied functions, each satisfying the constraint (12) for all and .
Similarly, the measure for is invariant under conjugation by any element in the stabilizer . So if acts transitively on , then it suffices to provide a single locally finite Borel measure that is invariant under conjugation by for some choice of basepoint . (In this case we set for each , which is not an overdetermined definition by our -invariance constraint on .)
This will be a recurring theme throughout the present paper. Oftentimes we will have some entity parametrized by in a -equivariant way entailing a -invariance (or later -equivariance) constraint for the single entity associated to any . The benefit of working with such parametrized entities is that it eliminates the need to check the independence of constructions from some choice of fundamental domain (or basepoint for transitive actions), which can be viewed as an implementation detail.
2.1.3 Generalization to Vector-Valued Functions
We also note that the cross-correlation defined in (10) readily generalizes to a transformation of vector-valued functions via matrix-valued filters. More specifically, for , for a continuous vector-valued function , and for a matrix-valued compactly supported filter satisfying the constraint (9) for all and , we obtain a cross-correlation in much the same way. In particular, the proof of -equivariance is provided by the same calculation (11). In the remainder of this Section 2, we generalize this form of a cross-correlation to -equivariant vector bundles.
2.2 Mackey Sections
First of all, recall from Section 2.1 that is a Hausdorff topological group and a Hausdorff -space. In order to reduce cross-correlations transforming sections of vector bundles to cross-correlations transforming vector-valued functions, Cohen et al. (2019, Section 2.3.1) proposed the use of Mackey functions. In this Section 2.2 we generalize or adapt this notion to the case of a not necessarily transitive -action.
To this end, let be a -equivariant real vector bundle over . Roughly speaking, this amounts to being a -space, the bundle projection being -equivariant, and each fiber above some (with respect to the projection ) carries the structure of a real vector space. A continuous map with for all is henceforth referred to as a continuous section of the bundle , and we denote the real vector space of such continuous sections by .
Now let be a continuous section. In order to transform by forming a cross-correlation, it will be convenient to express the value for some and as a vector in . To this end, we define the map
| (13) |
making the diagram
| (14) |
commute. Moreover, satisfies the two equations
| (15) | ||||
| (16) |
for all and .
Remark 2.1.
Definition 2.2 (Mackey Sections).
If acts transitively on , then (and hence ) is uniquely determined by the partially applied function for any choice of basepoint . Moreover, the equation (15) provides the -equivariance constraint
for any and . This is the type of function that Cohen et al. (2019, Section 2.3.1) refer to as a Mackey function, i.e. a function is a Mackey function (and hence determines a section of ) iff it satisfies such -equivariance constraint. So the Mackey section can also be thought of as the family of Mackey functions with respect to any basepoint in associated to the section .
2.2.1 Group Actions on (Mackey) Sections
We now define compatible -actions on sections and Mackey sections of the vector bundle . The associated -action on ordinary sections can be seen as a form of conjugation:
where
| (17) |
for all . As already noted above, in the present paper function application takes precedence over the binary operator “” denoting group actions. The corresponding -action on Mackey sections is more simple:
where
| (18) |
for all and .
Proposition 2.3.
The map
| (19) |
sending a section to its associated Mackey section is an -linear -equivariant bijection with inverse
| (20) |
Proof.
Let and let be the associated Mackey section. By the equation (16), the partial evaluation at the neutral element (20) is a left inverse to the map (19). As (15) is the defining equation of the Definition 2.2 of Mackey sections and as is completely determined by equations (15) and (16) considering Remark 2.1, the map (20) is a right inverse as well. For the proof of -equivariance let , and . Then we have
2.3 Cross-Correlations with a Filter
First of all, recall from Section 2.1 that
is a continuous family of locally finite Borel measures such that
| (8 revisited) |
for any and , where is the pushforward measure of along the conjugation
Now suppose we have -equivariant real vector bundles and . We aim to transform a section to a section of . So for each point we need to give a vector in in terms of . When doing this by a cross-correlation generalizing (10), we obtain such a vector in as a weighted sum or integral over the vector-values of the “Mackey function”
More specifically, the filter gives a linear map
| (21) |
for each and so a value in can be obtained as an integral
| (22) |
In order to formalize the idea that is continuous as an assignment of linear maps (21), we use the homomorphism bundle whose fiber above is the vector space of linear maps . So formally, we assume that is a continuous lift in the commutative diagram
| (23) |
Now in order for the integral (22) to be well-defined we impose that has compact support for any and in order for the map defined by to be -equivariant, we impose the constraint
| (24) |
for all , , and . Thus, if is some fundamental domain, then the filter is fully determined by all partially applied maps
for (where ). Moreover, for any one such partially applied map the equation (24) yields the constraint
for all , , and .
Now in light of Proposition 2.3, to describe a linear -equivariant map using we may as well describe a linear -equivariant map . This appears to be a sensible choice, considering the use of in the integral (22).
Definition 2.4 (Group Cross-Correlations).
For a Mackey section we define the cross-correlation by
| (25) |
for all and .
Theorem 2.5 (Well-Defined Cross-Correlations).
For the cross-correlation is indeed a Mackey section in the sense of Definition 2.2.
Proof.
By construction of we have the commutative diagram
Now let and . In order to complete our proof, we have to show the equation (15) with substituted for . And indeed we have
Here the fourth equality follows from the change of variables formula for the pushforward measure . ∎
Remark 2.6.
If the measure is left-invariant for all , then we may write the cross-correlation (25) also as
| (26) |
for all and , which is an adaptation of the formula provided by (Cohen et al., 2019, Equation 7) and (Gerken et al., 2023, Definition 3.8) to the case of a not necessarily transitive group action. By defining
we may then also write the cross-correlation (25) as a convolution:
for and . Note the subtle difference in notation with substituted for . Indeed, we have
However, as this can only be done for left-invariant measures, we stick with Definition 2.4 albeit the clumsy wording.
As we defined this general notion of a cross-correlation as a transformation outputting Mackey sections, the proof of -equivariance is a straightforward calculation.
Lemma 2.7 (Equivariance of Cross-Correlations).
The map
is -equivariant.
Proof.
For , , and we have
3 Orbitwise Integral Transforms
In the previous Section 2.3 we introduced a form of cross-correlation transforming sections of some -equivariant real vector bundle to sections of another such vector bundle . In the remainder of this paper, we compare this notion of a cross-correlation to that of an integral transform of sections for some kernel . Informally, a kernel is an assignment of a linear map
| (27) |
to any and any within the receptive field of ; so the value of the integral transform of a section at can be written as
| (28) |
with the domain of integration and its measure to be determined.
Now suppose is a section of and that is a filter as in the previous Section 2.3. Then the value of the resulting section at some point can be written as
| (22 revisited) |
Moreover, as the Mackey function only sees values of at points that are in the same orbit as , the receptive field of is constrained to its orbit . So in order to obtain an integral transform comparable to cross-correlations in the sense of Definition 2.4, we assume to be the domain of integration in (28) and the kernel to be defined on
where the right-hand side denotes the disjoint union as a set endowed with the subspace topology of . Then in order to use as a domain of integration in (28), we also need a locally finite Borel measure . So we also assume that we have a family
of locally finite Borel measures. As the scope of this paper is confined to -equivariant integral transforms, we further impose the equation
| (29) |
for all and , where is the pushforward measure of along the self-homeomorphism
In particular, the measure is -invariant for any . For an in depth discussion on how such families of measures satisfying even more restrictive constraints such as -invariance can be obtained in a natural way, consider the Appendix A.2.
Now in order to formalize the idea that is continuous as an assignment of linear maps (27), we view the natural surjections and as vector bundles over and we assume that is a continuous lift in the commutative diagram
with each partially applied lift for compactly supported. The associated orbitwise integral transform is then defined by
| (30) |
where denotes the vector space of all not necessarily continuous sections of the vector bundle . As we did not impose any relation or compatibility on the measures and for , we cannot guarantee continuity for sections obtained as the output of . However, if we can write using cross-correlations, as we will discuss in Section 4, then the continuity of the sections it outputs follows a posteriori.
3.1 Equivariance of Orbitwise Integral Transforms
In the context of -invariant measures on the base space , constraints on kernels entailing their integral transforms be -equivariant have been widely studied. In the following we show that essentially the same constraint as provided by (Gerken et al., 2023, Section 4.2) is sufficient and under suitable tameness assumptions also necessary for an orbitwise integral transform as defined by (30) to be -equivariant. More specifically, this constraint on a kernel as above is that we have the equation
| (31) |
for all , , , and .
Lemma 3.1.
If we have equation (31) for all , , , and , then the integral transform is -equivariant.
Proof.
For , and let . Then we have
Proposition 3.2.
Suppose that the integral transform is -equivariant. Moreover, for any we assume that the measure is strictly positive, that is paracompact, and that any compactly supported continuous section of the restricted vector bundle has a continuous extension to a section of , where . Then we have the equation (31) for all , , , and .
Proof.
Let , , and . Then we have
| (32) |
Now for let
By the previous equation (32) we have
| (33) |
for all sections and we have to show that for all and . To this end, it suffices to show that
for all linear forms , , and . Now let be a linear form and let be defined by
for all and , where is the restricted bundle of dual spaces. Moreover, suppose we have a compactly supported continuous section of the restricted vector bundle . By assumption there is a continuous extension of to a section defined on all of . Furthermore, we have
Thus, we conclude from Lemma B.1 that and hence
for all and . ∎
Corollary 3.3.
Suppose the action is transitive, that is paracompact, and that is strictly positive for some (hence any) . If is -equivariant, then we have the equation (31) for all , , , and .
4 Integral Transforms as Cross-Correlations
In this Section 4 we establish a close relationship between orbitwise integral transforms and cross-correlations. In particular, we will provide a construction for writing a -equivariant orbitwise integral transform associated to a kernel as a cross-correlation with a filter . As it turns out, such a filter may not be fully determined by . So in general, lifting an integral transform to a cross-correlation requires some choices to be made. Before we discuss this in full generality, we demonstrate at a simple example how this may require some trade-offs.
4.1 Example with Real Numbers and Integers
Let and let . We define a -action on by setting
for all and . Moreover, we assume we have identical trivial vector bundles
so we also identify their sections with continuous functions . Furthermore, the action on is confined to the first component:
for and . So for a function as a section of its associated Mackey section can be identified with the function
For the kernel provides a linear map , which we identify with the corresponding scalar coefficient making the kernel a continuous function . Similarly, we view any filter as a function .
Now let to be the Lebesgue measure and let be the product measure of the Lebesgue measure on and the counting measure on . As the Lebesgue measure also is a Haar measure and as such translation invariant, the constant family of measures satisfies the required constraint (29). Moreover, as is abelian the constant family of measures satisfies the constraint (8).
Now if we had a filter with the desired properties, then in particular the equation
| (34) |
would be satisfied for all continuous functions and all . So in order to find a filter satisfying (34), we may continuously choose for each pair an element with
| (35) |
and set
| (36) |
with all other values of set to . To this end, we could define the continuous map
| (37) |
which works for any kernel irrespective of its support.
4.1.1 Special Support of Kernel
In this case we may choose for the continuous map
| (38) |
As is abelian and as the action is transitive, we have by the constraint (24) for all , see also Section 2.1.2. Now let for some (hence any) . We consider the support of the filter depending on our choice for . If we use as specified in (37) to define using equation (36), then we have
| (39) |
whereas if we use as specified in (38), then we have
| (40) |
In case we have (40) for the support, then any discretization of can be represented by a fully populated 2D array, which is untrue for (39).
So clearly, there is a trade-off to be made here. On the one hand, we have the construction using (37) that works irrespective of the support of , and on the other hand there is the construction using (38) that only works for but it has the benefit that can be discretized by a fully populated 2D array. For this reason, our general construction of the filter will not only depend on the kernel itself but also on a choice as we had it here with .
4.1.2 Comparison to “Bi-Equivariant Kernels”
Before we proceed with the construction of filters from kernels, we use the above example, to compare the present notion of a group cross-correlation to the approach by Cohen et al. (2019), which is also surveyed in (Gerken et al., 2023, Section 3.2). In their terminology the function is a “one-argument kernel” and their constraint on (which in their notation is for both, their one- and their two-argument kernel) is bi-equivariance with respect to the stabilizer of the group action ; this is (Cohen et al., 2019, Theorem 3.2). Specialized to the present example of the abelian group and trivial vector bundles over , bi-equivariance amounts to invariance with respect to addition of elements in the stabilizer , i.e.
| (41) |
for all and . Now independent of our choice for the map , any filter obtained by the above construction will not satisfy this equation (41).
Now suppose that by some other means we obtained some continuous function satisfying the equation (41) for all and . As is an abelian group acting transitively on we now use the simplified form of a cross-correlation provided by (6) with respect to the measure . Moreover, let be a continuous function and let . Then as far as the first integral of the following calculation is finite, Fubini’s theorem implies
and hence . So for the group action , bi-equivariance of the filter/“one-argument kernel” results in vanishing, degenerate, or ill-defined cross-correlations. With that said, the present example is ruled out when assuming compact stabilizers as in (Gerken et al., 2023, Remark 3.2).
In summary, the constraint (24) on the filter , which specializes to no constraint on for the present example, is lenient enough to accommodate non-compact stabilizers. Moreover, while this additional flexibility also results in more than one filter providing the same integral transform, it allows the description as a cross-correlation to inform the shape of the tensor holding the trainable parameters of the filter (or in more general settings).
4.2 Compatible Measures
When defining cross-correlations in Definition 2.4, we used a family of measures defined on the group and for integral transforms we use a family of measures defined on the orbits of the action . In order to link these two families of measures, we assume there is a third family
of left-invariant locally finite Borel measures on the stabilizers of the action with the following two properties. Closely analogous to our constraints on the family we require that we have
for all and , where is the pushforward measure of along the conjugation
here as a map between stabilizers. Secondly, relating the three families of measures, we assume we have the equation
| (42) |
for all and compactly supported continuous functions . In this equation (42) we view as a pattern being matched against all within the domain of integration . More specifically, as we integrate over the free variable is bound to some such that . As the measure is left-invariant, the value of the inner integral
is independent of the particular choice of satisfying the equation .
Now in order to construct a filter from a kernel we will also need a way of associating to a real number a function (where ) whose integral evaluates to . If is compact, then we may define to be the constant function evaluating to . However, this construction only works when is finite and even then, we might prefer to concentrate the distribution of values towards the neutral element in order to limit the support of the resulting filter . To this end, we further assume we have a continuous function
where inherits the subspace topology from , with the following properties. For each the partially applied function is compactly supported and normed with respect to :
| (43) |
Moreover, we assume we have the constraint
| (44) |
for all , , and .
In most situations, we may well prefer to choose in such a way that integration of each partially applied function with over any Borel set of provides an approximation of the dirac measure on , which sends Borel sets containing the neutral element to and all other Borel sets to .
Examples 4.1.
In support of these assumptions, we provide the following two examples.
- 1.
In the example discussed in the previous Section 4.1, the stabilizer of any is the discrete additive subgroup
So we may choose to be the counting measure on for all . With and defined as in Section 4.1, we also have the equation (42) for all and compactly support continuous . Finally, we may define
which is easily seen to satisfy the equations (43) and (44). Using these choices with the general construction that follows in Section 4.4, we recover the filter we described in Section 4.1 (for the corresponding choice of ).
- 2.
For a more generic example, we assume the action satisfies the Assumption A.7. In this case, which is a generalization of (i), the Theorem A.9 provides families of measures , , and satisfying all of the above constraints (and more) as well as a function
with compactly supported for each , with
(62 anticipated) for all and , and with
(45) for all . Using the function we define the function
as its restriction to . As is compactly supported for each , corresponding partially applied functions are compactly supported as well. Moreover, the equations (43) and (44) follow directly from the equations (45) and (62) respectively.
4.3 Projection of Filters to Kernels
Before we get to lifting kernels to filters for cross-correlations, we provide a construction of the converse: How an integral transform can be obtained from a filter for cross-correlations. To this end, suppose we have -equivariant real vector bundles and and let be a filter as in Section 2.3. Then we define the kernel
by
| (46) |
for all , , and . As the Borel measure is left-invariant for any , the kernel is not overdetermined by these assignments (46).
Lemma 4.2.
Let , , , and . Then we have the equation
| (31 revisited) |
Proof.
Let with . Then we have
We also note that, as the map is continuous and as has compact support, the support of the partially applied map is compact as well for all .
Theorem 4.3.
For any continuous section and any we have
where is the Mackey section associated to in the sense of Definition 2.2.
Proof.
We have
4.4 Lifting Kernels to Filters
We now provide a converse to the construction of the previous Section 4.3. To this end, suppose we have -equivariant real vector bundles and and let
be a kernel as in Section 3. As a partially applied cross-correlation is -equivariant by Lemma 2.7 for a filter satisfying the constraint (24), the orbitwise integral transform is necessarily -equivariant if it can be written in terms of such cross-correlations. Moreover, under the mild tameness assumptions of Proposition 3.2, the integral transform is -equivariant iff its kernel satisfies the constraint
| (31 revisited) |
for all , , , and , which we assume to be satisfied from this point forward.
As we have seen already with the example of Section 4.1, lifting a kernel to a filter depends on a choice of the following. We need some closed subset that contains the support of and that is -invariant with respect to the diagonal action . Moreover, we need to choose some continuous map
subject to the constraint
| (35 revisited) |
for all . Now in order for our construction to promote the given constraint (31) for the kernel to the constraint (24), which we require from the filter , we need to impose an additional constraint on . More specifically, for any and we impose the equation
| (47) |
While we make no use of the so called category of elements associated to the group action (as a set-valued functor on ), it does provide the illustration
of this additional constraint (47).
Remark 4.4.
We add some comments related to the map .
- 1.
The additional constraint (47) is equivalent to the map being -equivariant with respect to the diagonal action on and the action by conjugation on .
- 2.
Let , , and let be the restriction of to . Then the partially applied map provides the trivialization
of the restricted vector bundle . As we require the support of the partially applied kernel to be contained in , the present construction can only be applied when the vector bundle is trivializable above the receptive field of . In order to circumvent this limitation, we generalize the present construction in Section 4.4.1.
As we collected all the necessary ingredients, we now define the filter by setting
| (48) |
for all , , and .
Lemma 4.5.
For the partially applied map has compact support.
Proof.
Let , let , let
and let
As is compact by our assumptions on the kernel and the function and as is continuous, it suffices to show the inclusion . To this end, let . Then we have and . As a result, we have and
Lemma 4.6.
Let , , and . Then we have the equation
| (24 revisited) |
Proof.
As the subset is assumed or chosen to be -invariant with respect to the diagonal action , we have iff we have . Thus, if , then we obtain the equation
and if , then we have
In view of the preceding Lemmas 4.5 and 4.6, the map can now be used as a filter in cross-correlations in the sense of Definition 2.4.
Theorem 4.7.
Let be a continuous section and let
be the corresponding Mackey section in the sense of Definition 2.2. Then we have
for any .
Proof.
Let , , and . For any we have
| (49) |
Hence, for any and we obtain the equation
| (50) |
As an end result we obtain
Corollary 4.8.
For any continuous section the section is continuous as well and hence a posteriori a vector in .
Proof.
Let . By the preceding Theorem 4.7 we have , which is continuous. ∎
In the case of a transitive action with compact stabilizers by a unimodular group , Aronsson (2022) has shown that any -equivariant transformation of vector bundle sections satisfying some additional tameness assumption can be obtained from cross-correlations. His construction is more abstract than the above construction of from and , and it is not clear whether the resulting filter (there referred to as a kernel and denoted by ) satisfies any particular constraints (such as bi-equivariance or a constraint similar to ours).
4.4.1 Generalization to Large Receptive Fields
The preceding construction lifting the filter to a filter depended on the choice of a -invariant subset supporting the kernel and a continuous map subject to the constraint
| (35 revisited) |
for all , which is also reasonable considering the example of Section 4.1. However, in view of Remark 4.4.(ii), the existence of such a map also entails that the vector bundle is trivializable above the receptive field of any . For most applications, this is likely no serious limitation, as the concept of a convolutional neural network is also informed by the idea that the receptive field of each node will be localized around this node. So it is not a far stretch to assume that the receptive field of each will be sufficiently local for to be trivializable above the receptive field of . With that said, it is still a reasonable question, whether the concept of a group convolutional layer can be pushed beyond vector bundles that are trivializable above each receptive field. And indeed, as with many other local constructions, such can be done using a sufficiently fine partition of unity.
We now assume there is a locally finite partition of unity
on that is -invariant with respect to the diagonal action . Moreover, let for and suppose we have a family
of continuous maps subject to the constraints
| (51) | ||||
| (52) |
for all , , and , where is the closure of in .
Also note, while our previous choice of a supporting region and a map had to take the support of the kernel into consideration, the partition of unity and the family of maps can be chosen independently from the kernel .
Now let and
| (53) |
As is -invariant with respect to the diagonal action , the constraint (31) on the kernel is inherited by this new kernel , i.e. we have
| (54) |
for all , , , and . Thus, we have the -equivariant integral transform
by Lemma 3.1. As we have by construction of and as we have the map as well as the constraints (54), (51), and (52), we are now back in the special case, where we can lift the kernel to a filter by setting
| (55) |
for all , , and in close analogy to (48).
Lemma 4.9.
For and the partially applied map has compact support.
Lemma 4.10.
Let , , , and . Then we have the equation
| (56) |
In view of the preceding Lemmas 4.9 and 4.10, the map can now be used as a filter in cross-correlations in the sense of Definition 2.4 for each .
Theorem 4.11.
Let , let be a continuous section, and let
be the corresponding Mackey section in the sense of Definition 2.2. Then we have
| (57) |
for any .
Proof.
This is Theorem 4.7 now just reformulated for the kernel , the integral transform , and the filter . ∎
With the preceding Theorem 4.11 we now managed to lift each localized kernel to a filter for . In order to lift the kernel itself, we intend to use the sum of the family of filters :
| (58) |
Note however, as the indexing set may be infinite, the sum (58) may not be well-defined a priori. In order to prove that (58) is well-defined and also in order to show that it does indeed lift the kernel , we now use the locally finite open cover of , to obtain a locally finite open cover of the orbit for each as follows. Let
for each and . Then the family is a locally finite open cover of for each .
Lemma 4.12.
For each the set
is finite.
Proof.
This follows from Lemma C.1 and our assumption that the partially applied kernel is compactly supported. ∎
Now let . For the partially applied filter vanishes iff the partially applied kernel vanishes. Thus, by the preceding Lemma 4.12 there are only finitely many with non-vanishing. As a result
is a finite sum of compactly supported maps , which implies the following.
Lemma 4.13.
For the partially applied map has compact support.
Lemma 4.14.
Let , , and . Then we have the equation
| (24 revisited) |
Proof.
We have
In view of the preceding Lemmas 4.13 and 4.14, the map can now be used as a filter in cross-correlations in the sense of Definition 2.4.
Theorem 4.15.
Let be a continuous section and let
be the corresponding Mackey section in the sense of Definition 2.2. Then we have
for any .
Proof.
Let . Then we have
Here the second but last equality follows from the family being a partition of unity. ∎
5 Other Integral Transforms
In the previous Section 4 we related cross-correlations to orbitwise integral transforms. Yet, in the literature more general integral transforms are considered as layers of group convolutional neural networks as well. In this Section 5 we discuss the trade-offs between group cross correlations in the sense of Definition 2.4 and more general integral transforms.
To this end, let and be Hausdorff -spaces with paracompact and let and be -equivariant vector bundles. Moreover, suppose we have a continuous compactly supported section of the homomorphism bundle as well as a -invariant locally finite strictly positive Borel measure . By reasoning similar to Section 3.1, we obtain that the associated integral transform
where
is -equivariant iff we have
| (59) |
for all , , , and ; see also (Gerken et al., 2023, Section 4.2).
For similar reasons as provided in Section 2.1.2, the kernel is fully determined by the family of partially applied maps
for some fundamental domain (and similarly for a fundamental domain of ). For each the partially applied map is then subject to the constraint
for all , , and . Now in some literature (in particular when is transitive so a single point suffices for a fundamental domain) an integral transform is also considered a convolution when expressed in terms of such partially applied maps; see for example (Gerken et al., 2023, Section 4.3).
There is a fundamental difference however. The partially applied map for each still is a section to the potentially twisted vector bundle . Whereas in the setting of Section 2.3, the partially applied map for some is just a vector-valued function. This is also illustrated by the diagram (23) where is pictured as a lift to a vector bundle over , while is a section to a vector bundle over the product . So in the generality considered in this paper, there are equivariant integral transforms that cannot be expressed in terms of cross correlations with the same benefits as Definition 2.4.
Funding
This research has been supported by EPSRC grant EP/Y028872/1, Mathematical Foundations of Intelligence: An “Erlangen Programme” for AI.
References
- Homogeneous vector bundles and g-equivariant convolutional neural networks. Sampl. Theory Signal Process. Data Anal. 20 (2) (en). Cited by: §4.4.
- Geometric Deep Learning: Grids, Groups, Graphs, Geodesics, and Gauges. Note: https://arxiv.org/abs/2104.13478 External Links: 2104.13478 Cited by: Appendix C.
- A general theory of equivariant cnns on homogeneous spaces. In Proceedings of the 33rd International Conference on Neural Information Processing Systems, Cited by: §1, §1, §2.1.2, §2.2, §2.2, Remark 2.6, §4.1.2, §4.1.2.
- Group equivariant convolutional networks. In Proceedings of The 33rd International Conference on Machine Learning, M. F. Balcan and K. Q. Weinberger (Eds.), Proceedings of Machine Learning Research, Vol. 48, New York, New York, USA, pp. 2990–2999. External Links: Link Cited by: §1.
- Geometric deep learning and equivariant neural networks. Artif. Intell. Rev. 56 (12), pp. 14605–14662 (en). Cited by: §2.1.2, Remark 2.6, §3.1, §4.1.2, §4.1.2, §5, §5.
- Vector bundles and k-theory. Note: https://pi.math.cornell.edu/~hatcher/VBKT/VBpage.html Cited by: Appendix B.
- On the generalization of equivariance and convolution in neural networks to the action of compact groups. In Proceedings of the 35th International Conference on Machine Learning, J. Dy and A. Krause (Eds.), Proceedings of Machine Learning Research, Vol. 80, pp. 2747–2755. External Links: Link Cited by: §1, §1, §2.1.2.
- Haar Measures. Note: https://arxiv.org/abs/2006.10956 External Links: 2006.10956 Cited by: §A.2, Assumption A.3.
Appendix A Constructing Families of Measures
In this Appendix A we show how one can construct families of measures as used in this paper in a natural way leaving as little to choice as possible. We start by providing the following notion.
Definition A.1.
We say that a -space is free, if there is a topological space and a -equivariant homeomorphism
Remark A.2.
Note that the -action associated to any free -space is necessarily fixed-point free. However, there are non-free -spaces with fixed-point free actions as for example as a -space under addition.
A.1 Families of Haar Measures
In order to provide a natural construction of a family of measures
we impose the following.
Assumption A.3.
We assume that is locally compact and that for any we have for the corresponding stabilizer, where is the kernel of the modular function , see for example (Tornier, 2020, Section 3). Moreover, we assume that has a -invariant locally finite partition of unity such that each induced open is a free -space in the sense of Definition A.1.
Examples A.4.
We provide the following two cases, where the Assumption A.3 is satisfied.
- 1.
If is unimodular, then and
So we may choose the single constant function
for a partition of unity .
- 2.
If acts transitively on , then acts freely and transitively on . So any choice of an -orbit yields an isomorphism
Thus, we may again choose the constant function
for a partition of unity .
Lemma A.5.
Under the Assumption A.3 there is a continuous function such that for all and .
Proof.
It suffices to provide a continuous function such that
for all and as is easily seen considering the diagram
To this end, let
be a locally finite -invariant partition of unity such that is a free -space for all . Moreover, we choose -equivariant homeomorphisms
| (60) |
for . By precomposing each of the functions
with the corresponding homeomorphism from (60) we obtain functions such that
| (61) |
for all , , and . Then we define
As for and and as the partition of unity is locally finite, the function is well-defined and continuous.
Now let and . Then we have
Here the third equality follows from the -invariance of the partition of unity and the sixth equality from
Theorem A.6.
Under the Assumption A.3 there is a continuous family
of Haar measures on such that
| (8 revisited) |
for all and , where is the pushforward measure of along the conjugation
Proof.
Let be some Haar measure on and let be as in the previous lemma. We set . Now let be some Borel set. Then we have
Finally, let be continuous and compactly supported. Then we have
Thus, the function
is continuous. ∎
A.2 Families of Group Invariant Measures
In addition to a family of Haar measures as in the previous Section A.1, we now describe the construction of compatible -invariant measures on the orbits of the action as well as Haar measures on the stabilizers.
Assumption A.7.
We impose the Assumption A.3 and moreover, we assume
- •
we have a non-vanishing compactly supported continuous function
invariant under conjugation by elements of
- •
and for any the stabilizer is unimodular and the map
is a quotient map.
Lemma A.8.
Under the Assumption A.7 there is a continuous function
such that
| (62) |
for all and and each partially applied function (where ) is a convex combination of functions of the form
with .
Proof.
It suffices to construct a continuous function
such that
for all and and each partially applied function (where ) is a convex combination of functions of the form
with . To this end, let
be a locally finite -invariant partition of unity such that is a free -space for all . Moreover, we choose -equivariant homeomorphisms
| (63) |
for . By combining the homeomorphisms (63) with the functions
we obtain functions
such that
and for all and and such that
| (64) |
for all , , and . Then we define
As for and and as the partition of unity is locally finite, the function is well-defined, continuous, and each partially applied function (where ) is a convex combination of functions of the form
with . Now let and . Then we have
Here the second equality follows from (64) and -invariance of the partition of unity . ∎
Theorem A.9.
Under the Assumption A.7 there is a family
of -invariant Radon measures, there are families
of Haar measures, and there is a function as in the previous Lemma A.8 such that
for all , such that
| (42 revisited) |
for all and compactly supported continuous , and such that
| (65) | ||||
for all and , where denotes the corresponding of the two vertical conjugation maps in
Moreover, the family is continuous and the partially applied function is compactly supported for each .
Remark A.10.
We note that as is stated to be -invariant for any in Theorem A.9, we could have stated equation (65) as instead. However, (65) is the equation that we will need and it is also more elementary to prove, i.e. without directly invoking -invariance.
Proof.
Let be as in the previous Lemma A.8 and suppose we have some . Let
be the unique Haar measures such that
It is well known there is a unique -invariant Radon measure
such that
| (42 revisited) |
for any compactly supported continuous function , see for example (Tornier, 2020, Theorem 4.2). Now let and . In order to show it suffices to test on the partially applied function as here
Completely analogously we obtain . Now let be continuous and compactly supported. Then we have
As we defined as the unique -invariant Radon measure satisfying (42) with substituted for , we get . Moreover, the partially applied function is compactly supported as it is a convex combination of functions of the form
with , and its integral is non-zero for the same reason.
Finally, let be continuous and compactly supported as before and let be some Haar measure on . Then we have
for all . As the integral is non-zero and continuous in , the integral is continuous in as well. ∎
Appendix B Vanishing Dual Sections
Let be a paracompact space, let be a real vector bundle over , and let be some strictly positive locally finite Borel measure on . Moreover, let be a compactly supported continuous section of the dual bundle associated to . (We use as a subscript to to denote compactly supported sections.)
Lemma B.1.
If we have
for all compactly supported continuous sections , then .
Proof.
By (Hatcher, 2003, Proposition 1.2) there is an inner product
on . Let be the section corresponding to under the musical isomorphism induced by the inner product . Then we have
for all and moreover,
hence for all . Applying the musical isomorphism once more we obtain . ∎
Appendix C Locally Finite Covers and Compact Subsets
Let be a topological space, let be a locally finite cover of , and let be a compact subset. (In our application of the following Lemma C.1, each subset is open in . In order to prove Lemma C.1, this assumption is not needed however.)
Lemma C.1.
The set is finite.
Proof.
For each we choose an open neighbourhood of such that for only a finite number of . Then we have the open cover of . Moreover, as is compact, there is a finite subset with
Now let . Then there is some
Thus, there is some with and hence
| (66) |