{"status":"ok","message-type":"work","message-version":"1.0.0","message":{"indexed":{"date-parts":[[2026,8,15]],"date-time":"2026-08-15T04:17:46Z","timestamp":1786767466999,"version":"3.56.0"},"reference-count":60,"publisher":"Association for Computing Machinery (ACM)","issue":"3","license":[{"start":{"date-parts":[[2019,8,8]],"date-time":"2019-08-08T00:00:00Z","timestamp":1565222400000},"content-version":"vor","delay-in-days":0,"URL":"https:\/\/2.zoppoz.workers.dev:443\/https\/www.acm.org\/publications\/policies\/copyright_policy#Background"}],"funder":[{"DOI":"10.13039\/501100001659","name":"Deutsche Forschungsgemeinschaft","doi-asserted-by":"publisher","award":["KO5206\/1-1, KR4661\/2-1"],"award-info":[{"award-number":["KO5206\/1-1, KR4661\/2-1"]}],"id":[{"id":"10.13039\/501100001659","id-type":"DOI","asserted-by":"publisher"}]}],"content-domain":{"domain":["dl.acm.org"],"crossmark-restriction":true},"short-container-title":["ACM Trans. Math. Softw."],"published-print":{"date-parts":[[2019,9,30]]},"abstract":"<jats:p>We present an algorithmic framework for matrix-free evaluation of discontinuous Galerkin finite element operators. It relies on fast quadrature with sum factorization on quadrilateral and hexahedral meshes, targeting general weak forms of linear and nonlinear partial differential equations. Different algorithms and data structures are compared in an in-depth performance analysis. The implementations of the local integrals are optimized by vectorization over several cells and faces and an even-odd decomposition of the one-dimensional interpolations. Up to 60% of the arithmetic peak on Intel Haswell, Broadwell, and Knights Landing processors is reached when running from caches and up to 40% of peak when also considering the access to vectors from main memory. On 2\u00d714 Broadwell cores, the throughput is up to 2.2 billion unknowns per second for the 3D Laplacian and up to 4 billion unknowns per second for the 3D advection on affine geometries, close to a simple copy operation at 4.7 billion unknowns per second. Our experiments show that MPI ghost exchange has a considerable impact on performance and we present strategies to mitigate this effect. Finally, various options for evaluating geometry terms and their performance are discussed. Our implementations are publicly available through the deal.II finite element library.<\/jats:p>","DOI":"10.1145\/3325864","type":"journal-article","created":{"date-parts":[[2019,8,8]],"date-time":"2019-08-08T12:30:31Z","timestamp":1565267431000},"page":"1-40","update-policy":"https:\/\/2.zoppoz.workers.dev:443\/https\/doi.org\/10.1145\/crossmark-policy","source":"Crossref","is-referenced-by-count":100,"title":["Fast Matrix-Free Evaluation of Discontinuous Galerkin Finite Element Operators"],"prefix":"10.1145","volume":"45","author":[{"ORCID":"https:\/\/2.zoppoz.workers.dev:443\/https\/orcid.org\/0000-0001-8406-835X","authenticated-orcid":false,"given":"Martin","family":"Kronbichler","sequence":"first","affiliation":[{"name":"Technical University of Munich, Germany"}],"role":[{"vocabulary":"crossref","role":"author"}]},{"given":"Katharina","family":"Kormann","sequence":"additional","affiliation":[{"name":"Max Planck Institute for Plasma Physics and Technical University of Munich, Germany"}],"role":[{"vocabulary":"crossref","role":"author"}]}],"member":"320","published-online":{"date-parts":[[2019,8,8]]},"reference":[{"key":"e_1_2_1_1_1","doi-asserted-by":"publisher","DOI":"10.1177\/1094342017694427"},{"key":"e_1_2_1_2_1","doi-asserted-by":"publisher","DOI":"10.1145\/3061708"},{"key":"e_1_2_1_3_1","doi-asserted-by":"publisher","DOI":"10.1515\/jnma-2018-0054"},{"key":"e_1_2_1_4_1","volume-title":"MFEM: Modular finite element methods. mfem.org.","author":"Anderson Robert","year":"2018","unstructured":"Robert Anderson , Andrew Barker , Jamie Bramwell , Jakub Cerveny , Johann Dahm , Veselin Dobrev , Yohann Dudouit , Aaron Fisher , Tzanio Kolev , Mark Stowell , and Vladimir Tomov . 2018 . MFEM: Modular finite element methods. mfem.org. Robert Anderson, Andrew Barker, Jamie Bramwell, Jakub Cerveny, Johann Dahm, Veselin Dobrev, Yohann Dudouit, Aaron Fisher, Tzanio Kolev, Mark Stowell, and Vladimir Tomov. 2018. MFEM: Modular finite element methods. mfem.org."},{"key":"e_1_2_1_5_1","doi-asserted-by":"publisher","DOI":"10.1137\/S0036142901384162"},{"key":"e_1_2_1_6_1","volume-title":"Karl Rupp, Barry F. Smith, Stefano Zampini, Hong Zhang, and Hong Zhang.","author":"Balay Satish","year":"2016","unstructured":"Satish Balay , Shrirang Abhyankar , Mark F. Adams , Jed Brown , Peter Brune , Kris Buschelman , Lisandro Dalcin , Victor Eijkhout , William D. Gropp , Dinesh Kaushik , Matthew G. Knepley , Lois Curfman McInnes , Karl Rupp, Barry F. Smith, Stefano Zampini, Hong Zhang, and Hong Zhang. 2016 . PETSc Users Manual. Technical Report ANL-95\/11 - Revision 3.7. Argonne National Laboratory . https:\/\/2.zoppoz.workers.dev:443\/http\/www.mcs.anl.gov\/petsc. Satish Balay, Shrirang Abhyankar, Mark F. Adams, Jed Brown, Peter Brune, Kris Buschelman, Lisandro Dalcin, Victor Eijkhout, William D. Gropp, Dinesh Kaushik, Matthew G. Knepley, Lois Curfman McInnes, Karl Rupp, Barry F. Smith, Stefano Zampini, Hong Zhang, and Hong Zhang. 2016. PETSc Users Manual. Technical Report ANL-95\/11 - Revision 3.7. Argonne National Laboratory. https:\/\/2.zoppoz.workers.dev:443\/http\/www.mcs.anl.gov\/petsc."},{"key":"e_1_2_1_7_1","doi-asserted-by":"publisher","DOI":"10.1145\/2049673.2049678"},{"key":"e_1_2_1_8_1","volume-title":"Software for Exascale Computing -- SPPEXA 2013-2015, Hans-Joachim Bungartz, Philipp Neumann, and Wolfgang E","author":"Bastian Peter","unstructured":"Peter Bastian , Christian Engwer , Jorrit Fahlke , Markus Geveler , Dominik G\u00f6ddeke , Oleg Iliev , Olaf Ippisch , Ren\u00e9 Milk , Jan Mohring , Steffen M\u00fcthing , Mario Ohlberger , Dirk Ribbrock , and Stefan Turek . 2016. Hardware-based efficiency advances in the EXA-DUNE project . In Software for Exascale Computing -- SPPEXA 2013-2015, Hans-Joachim Bungartz, Philipp Neumann, and Wolfgang E . Nagel (Eds.). Springer , Cham , 3--23. Peter Bastian, Christian Engwer, Jorrit Fahlke, Markus Geveler, Dominik G\u00f6ddeke, Oleg Iliev, Olaf Ippisch, Ren\u00e9 Milk, Jan Mohring, Steffen M\u00fcthing, Mario Ohlberger, Dirk Ribbrock, and Stefan Turek. 2016. Hardware-based efficiency advances in the EXA-DUNE project. In Software for Exascale Computing -- SPPEXA 2013-2015, Hans-Joachim Bungartz, Philipp Neumann, and Wolfgang E. Nagel (Eds.). Springer, Cham, 3--23."},{"key":"e_1_2_1_9_1","series-title":"Lecture Notes in Computer Science","volume-title":"Euro-Par 2014: Parallel Processing Workshops","author":"Bastian Peter","unstructured":"Peter Bastian , Christian Engwer , Dominik G\u00f6ddeke , Oleg Iliev , Olaf Ippisch , Mario Ohlberger , Stefan Turek , Jorrit Fahlke , Sven Kaulmann , Steffen M\u00fcthing , and Dirk Ribbrock . 2014. EXA-DUNE: Flexible PDE solvers, numerical methods and applications . In Euro-Par 2014: Parallel Processing Workshops . Lecture Notes in Computer Science , Vol. 8806 . Springer , 530--541. Peter Bastian, Christian Engwer, Dominik G\u00f6ddeke, Oleg Iliev, Olaf Ippisch, Mario Ohlberger, Stefan Turek, Jorrit Fahlke, Sven Kaulmann, Steffen M\u00fcthing, and Dirk Ribbrock. 2014. EXA-DUNE: Flexible PDE solvers, numerical methods and applications. In Euro-Par 2014: Parallel Processing Workshops. Lecture Notes in Computer Science, Vol. 8806. Springer, 530--541."},{"key":"e_1_2_1_10_1","doi-asserted-by":"publisher","DOI":"10.1007\/s10915-010-9396-8"},{"key":"e_1_2_1_11_1","doi-asserted-by":"publisher","DOI":"10.1016\/j.cpc.2015.02.008"},{"key":"e_1_2_1_12_1","doi-asserted-by":"publisher","DOI":"10.1007\/s10915-015-0049-9"},{"key":"e_1_2_1_13_1","volume-title":"Mund","author":"Deville Michel O.","year":"2002","unstructured":"Michel O. Deville , Paul F. Fischer , and Ernest H . Mund . 2002 . High-order Methods for Incompressible Fluid Flow. Vol. 9 . Cambridge University Press . Michel O. Deville, Paul F. Fischer, and Ernest H. Mund. 2002. High-order Methods for Incompressible Fluid Flow. Vol. 9. Cambridge University Press."},{"key":"e_1_2_1_14_1","volume-title":"Samuel D. Relton, Stanimire Tomov, and Mawussi Zounon.","author":"Dongarra Jack","year":"2016","unstructured":"Jack Dongarra , Iain Duff , Mark Gates , Azzam Haidar , Sven Hammarling , Nicholas J. Higham , Jonathan Hogg , Pedro Valero Lara , Samuel D. Relton, Stanimire Tomov, and Mawussi Zounon. 2016 . A Proposed API for Batched Basic Linear Algebra Subprograms. Technical Report. University of Tennessee . https:\/\/2.zoppoz.workers.dev:443\/https\/bit.ly\/batched-blas. Jack Dongarra, Iain Duff, Mark Gates, Azzam Haidar, Sven Hammarling, Nicholas J. Higham, Jonathan Hogg, Pedro Valero Lara, Samuel D. Relton, Stanimire Tomov, and Mawussi Zounon. 2016. A Proposed API for Batched Basic Linear Algebra Subprograms. Technical Report. University of Tennessee. https:\/\/2.zoppoz.workers.dev:443\/https\/bit.ly\/batched-blas."},{"key":"e_1_2_1_15_1","doi-asserted-by":"publisher","DOI":"10.1002\/fld.4511"},{"key":"e_1_2_1_16_1","doi-asserted-by":"publisher","DOI":"10.1002\/fld.4683"},{"key":"e_1_2_1_17_1","unstructured":"Paul Fischer Stefan Kerkemeier Adam Peplinski Dillon Shaver Ananias Tomboulides Misun Min Aleksandr Obabko and Elia Merzari. 2018. Nek5000 Web page. https:\/\/2.zoppoz.workers.dev:443\/https\/nek5000.mcs.anl.gov.  Paul Fischer Stefan Kerkemeier Adam Peplinski Dillon Shaver Ananias Tomboulides Misun Min Aleksandr Obabko and Elia Merzari. 2018. Nek5000 Web page. https:\/\/2.zoppoz.workers.dev:443\/https\/nek5000.mcs.anl.gov."},{"key":"e_1_2_1_18_1","volume-title":"Introduction to High Performance Computing for Scientists and Engineers","author":"Hager Georg","unstructured":"Georg Hager and Gerhard Wellein . 2011. Introduction to High Performance Computing for Scientists and Engineers . CRC Press , Boca Raton . Georg Hager and Gerhard Wellein. 2011. Introduction to High Performance Computing for Scientists and Engineers. CRC Press, Boca Raton."},{"key":"e_1_2_1_19_1","volume-title":"LIBXSMM: A high performance library for small matrix multiplications. https:\/\/2.zoppoz.workers.dev:443\/https\/github.com\/hfp\/libxsmm.","author":"Heinecke Alexander","year":"2017","unstructured":"Alexander Heinecke , Greg Henry , and Hans Pabst . 2017 . LIBXSMM: A high performance library for small matrix multiplications. https:\/\/2.zoppoz.workers.dev:443\/https\/github.com\/hfp\/libxsmm. Alexander Heinecke, Greg Henry, and Hans Pabst. 2017. LIBXSMM: A high performance library for small matrix multiplications. https:\/\/2.zoppoz.workers.dev:443\/https\/github.com\/hfp\/libxsmm."},{"key":"e_1_2_1_20_1","doi-asserted-by":"publisher","DOI":"10.1145\/1089014.1089021"},{"key":"e_1_2_1_21_1","volume-title":"Hesthaven and Tim Warburton","author":"Jan","year":"2008","unstructured":"Jan S. Hesthaven and Tim Warburton . 2008 . Nodal Discontinuous Galerkin Methods: Algorithms, Analysis , and Application. Texts in Applied Mathematics, Vol. 54 . Springer . Jan S. Hesthaven and Tim Warburton. 2008. Nodal Discontinuous Galerkin Methods: Algorithms, Analysis, and Application. Texts in Applied Mathematics, Vol. 54. Springer."},{"key":"e_1_2_1_22_1","doi-asserted-by":"publisher","DOI":"10.1016\/j.compfluid.2012.03.006"},{"key":"e_1_2_1_23_1","doi-asserted-by":"publisher","DOI":"10.1145\/2807591.2807644"},{"key":"e_1_2_1_24_1","unstructured":"M. Homolya R. C. Kirby and D. A. Ham. 2017. Exposing and exploiting structure: Optimal code generation for high-order finite element methods. arXiv preprint 1711.02473 (2017) cs.MS.  M. Homolya R. C. Kirby and D. A. Ham. 2017. Exposing and exploiting structure: Optimal code generation for high-order finite element methods. arXiv preprint 1711.02473 (2017) cs.MS."},{"key":"e_1_2_1_25_1","doi-asserted-by":"publisher","DOI":"10.1016\/j.jcp.2017.06.012"},{"key":"e_1_2_1_26_1","volume-title":"Intel 64 and IA-32 Architectures Optimization Reference Manual","author":"Intel Corporation 2017.","unstructured":"Intel Corporation 2017. Intel 64 and IA-32 Architectures Optimization Reference Manual . Intel Corporation . Order no. 248966-037, https:\/\/2.zoppoz.workers.dev:443\/https\/www.intel.com\/content\/dam\/www\/public\/us\/en\/documents\/manuals\/64-ia-32-architectures-optimization-manual.pdf. Intel Corporation 2017. Intel 64 and IA-32 Architectures Optimization Reference Manual. Intel Corporation. Order no. 248966-037, https:\/\/2.zoppoz.workers.dev:443\/https\/www.intel.com\/content\/dam\/www\/public\/us\/en\/documents\/manuals\/64-ia-32-architectures-optimization-manual.pdf."},{"key":"e_1_2_1_27_1","volume-title":"Knights Landing Edition. Morgan-Kaufmann","author":"Jeffers Jim","unstructured":"Jim Jeffers , James Reinders , and Avinash Sodani . 2016. Intel Xeon Phi Processor High Performance Programming , Knights Landing Edition. Morgan-Kaufmann , Cambridge , MA. Jim Jeffers, James Reinders, and Avinash Sodani. 2016. Intel Xeon Phi Processor High Performance Programming, Knights Landing Edition. Morgan-Kaufmann, Cambridge, MA."},{"key":"e_1_2_1_28_1","volume-title":"Sherwin","author":"Karniadakis George E.","year":"2005","unstructured":"George E. Karniadakis and Spencer J . Sherwin . 2005 . Spectral\/hp Element Methods for Computational Fluid Dynamics (2nd ed.). Oxford University Press . George E. Karniadakis and Spencer J. Sherwin. 2005. Spectral\/hp Element Methods for Computational Fluid Dynamics (2nd ed.). Oxford University Press."},{"key":"e_1_2_1_29_1","volume-title":"Automatic code generation for high-performance discontinuous Galerkin methods on modern architectures. arXiv preprint","author":"Kempf Dominic","year":"1812","unstructured":"Dominic Kempf , Ren\u00e9 Hess , Steffen M\u00fcthing , and Peter Bastian . 2018. Automatic code generation for high-performance discontinuous Galerkin methods on modern architectures. arXiv preprint 1812 .08075 (2018), math.NA. Dominic Kempf, Ren\u00e9 Hess, Steffen M\u00fcthing, and Peter Bastian. 2018. Automatic code generation for high-performance discontinuous Galerkin methods on modern architectures. arXiv preprint 1812.08075 (2018), math.NA."},{"key":"e_1_2_1_30_1","doi-asserted-by":"publisher","DOI":"10.1145\/2627373.2627387"},{"key":"e_1_2_1_31_1","doi-asserted-by":"publisher","DOI":"10.1016\/j.jcp.2009.06.041"},{"key":"e_1_2_1_32_1","volume-title":"Smith","author":"Knepley Matthew G.","year":"2013","unstructured":"Matthew G. Knepley , Jed Brown , Karl Rupp , and Barry F . Smith . 2013 . Achieving high performance with unified residual evaluation. arXiv preprint 1309.1204 (2013), cs.MS. Matthew G. Knepley, Jed Brown, Karl Rupp, and Barry F. Smith. 2013. Achieving high performance with unified residual evaluation. arXiv preprint 1309.1204 (2013), cs.MS."},{"key":"e_1_2_1_33_1","unstructured":"Dimitri Komatitsch Jean-Paul Ampuero Kangchen Bai Piero Basini C\u00e9line Blitz Ebru Bozdag Emanuele Casarotti Joseph Charles Min Chen Percy Galvez Dominik G\u00f6ddeke Vala Hj\u00f6rleifsd\u00f3ttir Sue Kientz Jes\u00fas Labarta Nicolas Le Goff Pieyre Le Loher Matthieu Lefebvre Qinya Liu Yang Luo Alessia Maggi Federica Magnoni Roland Martin Ren\u00e9 Matzen Dennis McRitchie Matthias Meschede Peter Messmer David Mich\u00e9a Surendra Nadh Somala Tarje Nissen-Meyer Daniel Peter Max Rietmann Elliott Sales de Andrade Brian Savage Bernhard Schuberth Anne Sieminski Leif Strand Carl Tape Jeroen Tromp Jean-Pierre Vilotte Zhinan Xie and Hejun Zhu. 2015. SPECFEM 3D Cartesian User Manual. Technical Report. Computational Infrastructure for Geodynamics Princeton University CNRS and University of Marseille and ETH Z\u00fcrich.  Dimitri Komatitsch Jean-Paul Ampuero Kangchen Bai Piero Basini C\u00e9line Blitz Ebru Bozdag Emanuele Casarotti Joseph Charles Min Chen Percy Galvez Dominik G\u00f6ddeke Vala Hj\u00f6rleifsd\u00f3ttir Sue Kientz Jes\u00fas Labarta Nicolas Le Goff Pieyre Le Loher Matthieu Lefebvre Qinya Liu Yang Luo Alessia Maggi Federica Magnoni Roland Martin Ren\u00e9 Matzen Dennis McRitchie Matthias Meschede Peter Messmer David Mich\u00e9a Surendra Nadh Somala Tarje Nissen-Meyer Daniel Peter Max Rietmann Elliott Sales de Andrade Brian Savage Bernhard Schuberth Anne Sieminski Leif Strand Carl Tape Jeroen Tromp Jean-Pierre Vilotte Zhinan Xie and Hejun Zhu. 2015. SPECFEM 3D Cartesian User Manual. Technical Report. Computational Infrastructure for Geodynamics Princeton University CNRS and University of Marseille and ETH Z\u00fcrich."},{"key":"e_1_2_1_34_1","volume-title":"Implementing Spectral Methods for Partial Differential Equations","author":"Kopriva David","unstructured":"David Kopriva . 2009. Implementing Spectral Methods for Partial Differential Equations . Springer , Berlin . David Kopriva. 2009. Implementing Spectral Methods for Partial Differential Equations. Springer, Berlin."},{"key":"e_1_2_1_35_1","doi-asserted-by":"publisher","DOI":"10.4208\/cicp.101214.021015a"},{"key":"e_1_2_1_36_1","doi-asserted-by":"publisher","DOI":"10.1109\/eScience.2011.53"},{"key":"e_1_2_1_37_1","doi-asserted-by":"publisher","DOI":"10.1016\/j.jcp.2017.07.039"},{"key":"e_1_2_1_38_1","doi-asserted-by":"crossref","unstructured":"Martin Kronbichler and Momme Allalen. 2018. Efficient high-order discontinuous Galerkin finite elements with matrix-free implementations. In Advances and Trends in Environmental Informatics H.-J. Bungartz D. Kranzlm\u00fcller V. Weinberg J. Weism\u00fcller and V. Wohlgemuth (Eds.). 89--110.  Martin Kronbichler and Momme Allalen. 2018. Efficient high-order discontinuous Galerkin finite elements with matrix-free implementations. In Advances and Trends in Environmental Informatics H.-J. Bungartz D. Kranzlm\u00fcller V. Weinberg J. Weism\u00fcller and V. Wohlgemuth (Eds.). 89--110.","DOI":"10.1007\/978-3-319-99654-7_7"},{"key":"e_1_2_1_39_1","doi-asserted-by":"publisher","DOI":"10.5555\/3195466.3195472"},{"key":"e_1_2_1_40_1","doi-asserted-by":"publisher","DOI":"10.1016\/j.compfluid.2012.04.012"},{"key":"e_1_2_1_41_1","doi-asserted-by":"publisher","DOI":"10.1007\/978-3-319-58667-0_13"},{"key":"e_1_2_1_42_1","doi-asserted-by":"publisher","DOI":"10.1002\/nme.5137"},{"key":"e_1_2_1_43_1","doi-asserted-by":"publisher","DOI":"10.1137\/16M110455X"},{"key":"e_1_2_1_44_1","doi-asserted-by":"publisher","DOI":"10.1145\/3054944"},{"key":"e_1_2_1_45_1","doi-asserted-by":"publisher","DOI":"10.1109\/SC.2014.28"},{"key":"e_1_2_1_46_1","doi-asserted-by":"publisher","DOI":"10.1137\/15M1021167"},{"key":"e_1_2_1_47_1","doi-asserted-by":"publisher","DOI":"10.1016\/j.cageo.2016.03.008"},{"key":"e_1_2_1_48_1","volume-title":"High-performance implementation of matrix-free high-order discontinuous Galerkin methods. arXiv preprint 1711.10885","author":"M\u00fcthing Steffen","year":"2017","unstructured":"Steffen M\u00fcthing , Marian Piatkowski , and Peter Bastian . 2017. High-performance implementation of matrix-free high-order discontinuous Galerkin methods. arXiv preprint 1711.10885 ( 2017 ), math.NA. Steffen M\u00fcthing, Marian Piatkowski, and Peter Bastian. 2017. High-performance implementation of matrix-free high-order discontinuous Galerkin methods. arXiv preprint 1711.10885 (2017), math.NA."},{"key":"e_1_2_1_49_1","doi-asserted-by":"publisher","DOI":"10.1016\/0021-9991(80)90005-4"},{"key":"e_1_2_1_50_1","doi-asserted-by":"publisher","DOI":"10.1016\/0021-9991(84)90128-1"},{"key":"e_1_2_1_51_1","doi-asserted-by":"publisher","DOI":"10.1145\/2998441"},{"key":"e_1_2_1_52_1","unstructured":"James Reinders. 2007. Intel Threading Building Blocks. O\u2019Reilly.   James Reinders. 2007. Intel Threading Building Blocks. O\u2019Reilly."},{"key":"e_1_2_1_53_1","doi-asserted-by":"publisher","DOI":"10.1016\/j.jcp.2016.08.005"},{"key":"e_1_2_1_54_1","volume-title":"Technical Report ASC Report No. 30\/2014","author":"Sch\u00f6berl Joachim","unstructured":"Joachim Sch\u00f6berl . 2014. C++11 Implementation of Finite Elements in NG Solve . Technical Report ASC Report No. 30\/2014 . Vienna University of Technology . Joachim Sch\u00f6berl. 2014. C++11 Implementation of Finite Elements in NGSolve. Technical Report ASC Report No. 30\/2014. Vienna University of Technology."},{"key":"e_1_2_1_55_1","doi-asserted-by":"publisher","DOI":"10.1137\/18M1185399"},{"key":"e_1_2_1_56_1","doi-asserted-by":"publisher","DOI":"10.1006\/jcph.1996.0042"},{"key":"e_1_2_1_57_1","volume-title":"Kelly","author":"Sun Tianjiao","year":"2019","unstructured":"Tianjiao Sun , Lawrence Mitchell , Kaushik Kulkarni , Andreas Kl\u00f6ckner , David A. Ham , and Paul H. J . Kelly . 2019 . A study of vectorization for matrix-free finite element methods. arXiv preprint 1903.08243 (2019), cs.MS. Tianjiao Sun, Lawrence Mitchell, Kaushik Kulkarni, Andreas Kl\u00f6ckner, David A. Ham, and Paul H. J. Kelly. 2019. A study of vectorization for matrix-free finite element methods. arXiv preprint 1903.08243 (2019), cs.MS."},{"key":"e_1_2_1_58_1","doi-asserted-by":"publisher","DOI":"10.1109\/ICPPW.2010.38"},{"key":"e_1_2_1_59_1","doi-asserted-by":"publisher","DOI":"10.1002\/fld.3767"},{"key":"e_1_2_1_60_1","doi-asserted-by":"publisher","DOI":"10.1145\/1498765.1498785"}],"container-title":["ACM Transactions on Mathematical Software"],"original-title":[],"language":"en","link":[{"URL":"https:\/\/2.zoppoz.workers.dev:443\/https\/dl.acm.org\/doi\/10.1145\/3325864","content-type":"unspecified","content-version":"vor","intended-application":"text-mining"},{"URL":"https:\/\/2.zoppoz.workers.dev:443\/https\/dl.acm.org\/doi\/pdf\/10.1145\/3325864","content-type":"unspecified","content-version":"vor","intended-application":"similarity-checking"}],"deposited":{"date-parts":[[2025,6,17]],"date-time":"2025-06-17T23:53:08Z","timestamp":1750204388000},"score":1,"resource":{"primary":{"URL":"https:\/\/2.zoppoz.workers.dev:443\/https\/dl.acm.org\/doi\/10.1145\/3325864"}},"subtitle":[],"short-title":[],"issued":{"date-parts":[[2019,8,8]]},"references-count":60,"journal-issue":{"issue":"3","published-print":{"date-parts":[[2019,9,30]]}},"alternative-id":["10.1145\/3325864"],"URL":"https:\/\/2.zoppoz.workers.dev:443\/https\/doi.org\/10.1145\/3325864","relation":{},"ISSN":["0098-3500","1557-7295"],"issn-type":[{"value":"0098-3500","type":"print"},{"value":"1557-7295","type":"electronic"}],"subject":[],"published":{"date-parts":[[2019,8,8]]},"assertion":[{"value":"2017-11-01","order":0,"name":"received","label":"Received","group":{"name":"publication_history","label":"Publication History"}},{"value":"2019-04-01","order":1,"name":"accepted","label":"Accepted","group":{"name":"publication_history","label":"Publication History"}},{"value":"2019-08-08","order":2,"name":"published","label":"Published","group":{"name":"publication_history","label":"Publication History"}}]}}