Distribution-aware Block-sparse Recovery via Convex Optimization
Abstract
We study the problem of reconstructing a block-sparse signal from compressively sampled measurements. In certain applications, in addition to the inherent block-sparse structure of the signal, some prior information about the block support, i.e. blocks containing non-zero elements, might be available. Although many block-sparse recovery algorithms have been investigated in Bayesian framework, it is still unclear how to incorporate the information about the probability of occurrence into regularization-based block-sparse recovery in an optimal sense. In this work, we bridge between these fields by the aid of a new concept in conic integral geometry. Specifically, we solve a weighted optimization problem when the prior distribution about the block support is available. Moreover, we obtain the unique weights that minimize the expected required number of measurements. Our simulations on both synthetic and real data confirm that these weights considerably decrease the required sample complexity.
Index Terms:
Block sparse recovery, Bayesian information, Conic integral geometry, Convex optimization.I Introduction
Compressed Sensing (CS) has emerged in the past decade as a modern technique for recovering a sparse vector from compressed measurements (see [1, 2] for more explanations about this field). In this work, we consider signals that have a block-sparse structure, namely their non-zero entries appear in blocks. This property has been referred in the literature as block-sparsity. It is common to use the following optimization problem for recovering the signal from compressive measurements.
| (1) |
where represents a fat measurement matrix with , are the default disjoint blocks of size that partition the set , is the observation vector11 1 Our analysis holds for both real and complex-valued signals and measurements., is the noise term which is considered to be i.i.d. Gaussian with variance , and is an upper-bound for . Most of the earlier literature in block-sparse recovery is focused on the case of single constraint . However, in many applications such as DNA micro-arrays [3, 4], computational neuroscience [5] multi-band signal reconstruction, multiple measurement vector (MMV) problem [6], and the reconstruction of signals in union of subspaces [7, 8] [9, 10], there exist additional information (or alternatively additional constraints in ) about the signal of interest. In this work, we explore the benefits of having access to extra information about the distribution of the block support (blocks containing non-zero elements) on the required number of measurements. To this end, we propose the optimization problem
where the quantities are some positive scalars, and is some predefined model that restricts the feasible set of the solution. Specifically, we consider two new models for prior information that are of practical interest:
- •
Model 1 (Prior distribution) : We assume that the prior distribution of the block support is available. Under this setting, there are known probabilities associated with each block index . Namely,
(2) where returns the block support of a vector.
- •
Model 2 (Multiple block support estimates): We consider disjoint sets with that intersect with the expected accuracy
(3) To each subset , we assign a fixed weight . In fact, it holds that
(4) where is the indicator function of the set , , and .
One of the applications of Model 1 and 2 is in direction of arrival (DOA) estimation. In this application, Model 1 implies that the probability of having a target in an angle indexed by is known in advance [11]. In Model 2, however, we assume to know the expected number of targets in a range of angles represented by . Such statistics might be available from previous measurements in a dynamic scenario. Obviously, Model 1 imposes more strict conditions as knowing the probabilities for all angles is not easily achievable. In contrast, Model 2 could be applicable as the full angular range could be divided into 3 or 4 intervals, for which we can evaluate the average number of included targets.
In this work, we obtain the weights and that minimize a threshold describing the expected number of required measurements for Models 1 and 2, respectively. Our approach is to find a suitable upper-bound for . The bound is not necessarily tight but leads to closed-form expressions for and in Models 1 and 2, respectively.
I-A Related works
CS in presence of prior information has been studied in different signal models. While a large part of research (see for example [12, 13, 14, 15, 16, 17]) deals with deterministic signal models, only a few works (see [18, 19, 20, 21]) have investigated random signal models with Bayesian information. In the deterministic model, the ground-truth signal has intersected with a few sets which called support estimates. The contributing level of each set to the support is available to the experimenter[13, 12]. This exact situation is investigated in [17]. They propose a non-uniform model for capturing deterministic prior information. The work [19] considers a probabilistic model where there is a continuous shape function describing the probability of contributing each element to the support. The authors obtain an upper-bound for failure probability of weighted minimization. Their approach is based on calculating the internal and external angles of a weighted cross polytope. With a different approach, [20] has investigated a discrete measure for Bayesian information (a special case of Model 1 with ). They relate the weights of weighted minimization to the discrete probability distribution by minimizing the expected intrinsic volumes of a weighted cone.
I-B Contributions
As listed below, we have three main contributions in this work. Besides, our results are also applicable in DOA estimation (see Section I-C) and functional magnetic resonance imaging (fMRI) reconstruction with parallel coils [22].
- 1.
Optimally exploiting the block distribution. In presence of a block distribution, we obtain the optimal weights in . This result can be considered as an extension of [20] to the block-sparse (and joint-sparse22 2 In this case, the non-zero blocks have common support.) setting. However, the derivation of the optimal weights in this case is non-trivial and rather challenging. Further, our derivation approach is different from [20].
- 2.
Optimally exploiting multiple estimates. In presence of multiple block-support estimates, we derive the optimal penalizing coefficients corresponding to each set in weighted and minimizations.
- 3.
Robustness of optimal weights against inaccurate prior information. We analytically examine how close one can get to the optimal weights if the prior information (s and s) is inaccurate. This result is important in practical scenarios, as we only have access to the approximations of s and s.
I-C Application (Broadband DOA Estimation)
Suppose that far-field broadband signals in the frequency range incident on an -element uniform linear array (ULA). The received signal in sensors at time and -th frequency bin can be expressed as:
| (5) |
where is the steering vector, is the propagation velocity, is the inter-sensor spacing and is a Gaussian noise term with variance . In practice, one has to take several snapshots . This temporal redundancy is crucial in practice since the array size is limited due to physical constraints [23, 24]. Consequently, one may write
| (6) |
where , and is defined similar to . If the sources is time-invariant over the period of snapshotting, then for all the non-zero dominant peaks in occur at the same locations corresponding to the ground-truth DOAs. Hence, DOA estimation can be cast as recovering a joint sparse signal from . In addition, it is realistic for a radar engineer to know the probability of appearing the ground-truth DOAs in some angular bands [11] (see Figure 1 for a schematic model of this scenario).
Notation. Throughout, scalars are denoted by lowercase letters, vectors by lowercase boldface letters, and matrices by uppercase boldface letters. The th element of a vector is shown either by or . denotes the polar of a cone . We show sets (e.g. ) by calligraphic uppercase letters. We show the set by . is used to represent the complement of a set .
II Main results
In the following propositions, we obtain closed-form solutions for optimal weights in case that satisfies Model 1 and 2. The proofs are provided in Appendix -B.
Proposition 1.
Let satisfy Model 1 with parameter . Then, the optimal weights in are obtained by solving the following equations simultaneously:
| (7) |
Remark 1.
(Prior work) The special case ( is sparse instead of block-sparse) reduces (1) to the weighted minimization which is studied in [20]. Therefore, Proposition 1 generalizes the results of [20] to the block-sparse case. However, our approach to reach this generalized result is different from and somewhat simpler than [20].
Proposition 2.
Let be decomposed into blocks of equal size . Assume that there exist independent estimates of with parameter . Then, the optimal weights in Model 2 are obtained by solving
| (8) |
Remark 2.
The optimal weights s and s in Propositions 1 and 2 are respectively obtained by minimizing an upper-bound of the expected number of required measurements in problems and . We numerically observe that the upper-bound is tight for non-uniform distributions of , but the exact identification of such distributions is beyond the scope of this work. Further, the system of equations in (1) and (2) are solved using the function in MATLAB.
In applications, it is of important practical value to know how the inaccuracies of and in Model 1 and 2, respectively, affect the optimal weights. The following theorem is about this challenge.
Theorem 1.
Assume that and be the true and approximate estimate of , respectively. Let and be the optimal weights corresponding to , and . Then, there exists a constant such that
| (9) |
where is incomplete gamma function and is the nonlinear function in (1).
From the above theorem and the right image of Figure 2, one can infer that the method of obtaining and is robust to slight changes of and as long as they are greater than approximately .
III Simulations
In the first experiment, we construct a random block-sparse , whose building blocks have equal size . The probability of activating each block is taken from the vector depicted in the left image of Figure 2 . Then, this signal is observed through a measurement matrix , of which the elements are drawn from i.i.d. standard normal distribution. We obtain the optimal weights corresponding to by solving (1). We also examine the heuristic weights . In the middle image of Figure 2, we plot the success rate as a function of . For each , we average over realizations of and . As expected, requires fewer measurements for successful recovery with optimal weights rather than the other two options (heuristic or equal weights). In turn, heuristic weights are also superior to the equal weights.
In the second experiment, we test DOA estimation using broadband signals (see Subsection I-C and Figure 1). The angular half-space is divided into angular grids. For each frequency bin GHz, we take snapshots. We assume that there exist three () sets with expected accuracies , , and . Also, we use sensors for recovering ground-truth sources (located at the angles , , , , , , , , , ) and implement the optimization problem when is chosen optimally (i.e. ’s are obtained using (2)), heuristically (i.e. ) and equally. As it turns out from Figure 3, while with equal and heuristic weights detects [respectively many and a few] non-existing sources with spurious DOAs, locates the ground-truth sources correctly. This in turn suggests that our optimal weighting strategy considerably decreases the required number of sensors.
-A Preliminaries
In this section, we introduce two concepts from conic integral geometry that are used in our analysis.
Descent cone: Let be a vector with a special low-dimensional structure (e.g. block-sparsity). Assume that is a convex function that promotes this structure. Then, the set of decent directions forms a convex set defined by:
| (10) |
Statistical dimension: Statistical dimension is intuitively a measure for the size of a cone. It is shown in [25] that statistical dimension of the above decent cone defined by
| (11) |
specifies the required number of measurements (i.e. ) that
| (12) |
needs for perfect recovery. Here, is a vector with i.i.d. standard normal distribution. We define the expected number of measurements needed for as
| (13) |
Then, we call the weights that minimize optimal in sense of expected sample complexity.
-B Proof of Propositions 1 and 2
Proof.
Since is upper-bounded by an expression that only depends on and (see Lemma 1), we show the corresponding upper-bound by . It holds that:
| (14) |
We proceed by using a closed-form expression for a special case of which is obtained in [17, Lemma 2].
Lemma 1.
The statistical dimension of descent cone of any vector with satisfies:
Moreover, the minimum is achieved at a unique . The inequality is in fact equality in the asymptotic case ()[17, Proposition 3].
Thus, by the aid of Lemma 1, it holds that
| (15) |
where in the inequality , the notation signifies the indicator function of an event . The inequality results from the Jensen inequality for concave functions. Therefore, we have:
| (16) |
Moreover, by partitioning the set into the sets s, one may write:
| (17) |
where we used the relation (4) in the last equality. By further assumption , it holds that:
| (18) |
Since and , we have:
| (19) |
As multiplication with a positive scalar keeps the optimal choices of and in and untouched, by minimizing the expressions in brackets in (16) and (19) with respect to and (the second derivatives of these expressions are always positive with respect to and and hence they are strictly convex functions), one can get to the expressions (1) and (2). ∎
References
- [1] E. J. Candes and T. Tao, “Decoding by linear programming,” IEEE transactions on information theory, vol. 51, no. 12, pp. 4203–4215, 2005.
- [2] D. L. Donoho, “For most large underdetermined systems of linear equations the minimal -norm solution is also the sparsest solution,” Communications on pure and applied mathematics, vol. 59, no. 6, pp. 797–829, 2006.
- [3] F. Parvaresh, H. Vikalo, S. Misra, and B. Hassibi, “Recovering sparse signals using sparse measurement matrices in compressed dna microarrays,” IEEE Journal of Selected Topics in Signal Processing, vol. 2, no. 3, pp. 275–285, 2008.
- [4] M. Stojnic, F. Parvaresh, and B. Hassibi, “On the reconstruction of block-sparse signals with an optimal number of measurements,” arXiv preprint arXiv:0804.0041, 2008.
- [5] T. Euler and T. Baden, “Computational neuroscience: Species-specific motion detectors,” Nature, 2016.
- [6] M. Mishali and Y. C. Eldar, “Reduce and boost: Recovering arbitrary sets of jointly sparse vectors,” IEEE Transactions on Signal Processing, vol. 56, no. 10, pp. 4692–4702, 2008.
- [7] Y. C. Eldar and M. Mishali, “Robust recovery of signals from a structured union of subspaces,” IEEE Transactions on Information Theory, vol. 55, no. 11, pp. 5302–5316, 2009.
- [8] Y. M. Lu and M. N. Do, “A theory for sampling signals from a union of subspaces,” IEEE transactions on signal processing, vol. 56, no. 6, pp. 2334–2345, 2008.
- [9] M. Mishali and Y. C. Eldar, “Blind multiband signal reconstruction: Compressed sensing for analog signals,” IEEE Transactions on Signal Processing, vol. 57, no. 3, pp. 993–1009, 2009.
- [10] M. Mishali and Y. C. Eldar, “From theory to practice: Sub-nyquist sampling of sparse wideband analog signals,” IEEE Journal of Selected Topics in Signal Processing, vol. 4, no. 2, pp. 375–391, 2010.
- [11] K. V. Mishra, M. Cho, A. Kruger, and W. Xu, “Spectral super-resolution with prior knowledge,” IEEE transactions on signal processing, vol. 63, no. 20, pp. 5342–5357, 2015.
- [12] D. Needell, R. Saab, and T. Woolf, “Weighted-minimization for sparse recovery under arbitrary prior information,” Information and Inference: A Journal of the IMA, vol. 6, no. 3, pp. 284–309, 2017.
- [13] M. A. Khajehnejad, W. Xu, A. S. Avestimehr, and B. Hassibi, “Analyzing weighted minimization for sparse recovery with nonuniform sparse models,” IEEE Transactions on Signal Processing, vol. 59, no. 5, pp. 1985–2001, 2011.
- [14] N. Vaswani and W. Lu, “Modified-cs: Modifying compressive sensing for problems with partially known support,” IEEE Transactions on Signal Processing, vol. 58, no. 9, pp. 4595–4607, 2010.
- [15] R. G. Baraniuk, V. Cevher, M. F. Duarte, and C. Hegde, “Model-based compressive sensing,” IEEE Transactions on Information Theory, vol. 56, no. 4, pp. 1982–2001, 2010.
- [16] S. Oymak, M. A. Khajehnejad, and B. Hassibi, “Recovery threshold for optimal weight minimization,” in Information Theory Proceedings (ISIT), 2012 IEEE International Symposium on, pp. 2032–2036, IEEE, 2012.
- [17] S. Daei, F. Haddadi, and A. Amini, “Exploiting prior information in block sparse signals,” arXiv preprint arXiv:1804.08444, 2018.
- [18] W. Xu, Compressive sensing for sparse approximations: constructions, algorithms, and analysis. PhD thesis, California Institute of Technology, 2010.
- [19] S. Misra and P. A. Parrilo, “Weighted -minimization for generalized non-uniform sparse model,” IEEE Transactions on Information Theory, vol. 61, no. 8, pp. 4424–4439, 2015.
- [20] M. Díaz, M. Junca, F. Rincón, and M. Velasco, “Compressed sensing of data with a known distribution,” Applied and Computational Harmonic Analysis, 2017.
- [21] J. Fang, Y. Shen, H. Li, and P. Wang, “Pattern-coupled sparse bayesian learning for recovery of block-sparse signals,” IEEE Transactions on Signal Processing, vol. 63, no. 2, pp. 360–372, 2015.
- [22] H. Jung, K. Sung, K. S. Nayak, E. Y. Kim, and J. C. Ye, “k-t focuss: a general compressed sensing framework for high resolution dynamic mri,” Magnetic resonance in medicine, vol. 61, no. 1, pp. 103–116, 2009.
- [23] Z. Yang and L. Xie, “Exact joint sparse frequency recovery via optimization methods,” IEEE Transactions on Signal Processing, vol. 64, no. 19, pp. 5145–5157, 2016.
- [24] M. M. Hyder and K. Mahata, “Direction-of-arrival estimation using a mixed norm approximation,” IEEE Transactions on Signal Processing, vol. 58, no. 9, pp. 4646–4655, 2010.
- [25] D. Amelunxen, M. Lotz, M. B. McCoy, and J. A. Tropp, “Living on the edge: Phase transitions in convex programs with random data,” Information and Inference: A Journal of the IMA, vol. 3, no. 3, pp. 224–294, 2014.