Data Migration

Explore top LinkedIn content from expert professionals.

  • View profile for Josh Aharonoff, CPA

    Building World-Class Financial Models in Minutes | 485K+ Followers | Founder @ Mighty Digits

    485,883 followers

    Old Excel vs Modern Excel Which one are you using? Excel is a completely different beast than what we used 5 years ago. But most of us still work in it like it's 2015. Let's fix that - here are the powerful features you might be missing 👇 💾 OLD WAY → Reference Functions VLOOKUP ruled our spreadsheets for years, but it came with frustrating limitations. You could only search left to right, vertical searches were impossible, and error handling? Almost non-existent. Adding a column meant watching your formulas break in real-time. 💻 NEW WAY → XLOOKUP XLOOKUP changed everything. Now you can: ✅ Search in any direction you want ✅ Use built-in approximate matching ✅ Handle errors intelligently ✅ Process large datasets at lightning speed No more complex nested formulas just to look up a simple value. 💾 OLD WAY → Basic PivotTables Remember when PivotTables could only handle basic summarization? One table at a time, simple calculations only, and those dreaded memory constraints that popped up with any decent-sized dataset. 💻 NEW WAY → Power Pivot Power Pivot transforms how we analyze data: ✅ Connect multiple tables with real relationships ✅ Use DAX formulas for complex calculations ✅ Process millions of rows without breaking a sweat ✅ Built-in time intelligence functions that save hours Your reports just got a serious upgrade. 💾 OLD WAY → Manual Data Cleanup The old way meant endless hours of: - Copy-paste operations that numbed your brain - Find and replace marathons - Manual data cleaning that took forever - Constant risk of human error 💻 NEW WAY → Power Query Power Query is like having a data cleaning assistant: ✅ Connects multiple data sources automatically ✅ Transforms data with reusable steps ✅ Cleans data with advanced rules ✅ Refreshes with one click Set it up once, use it forever. That's work smarter, not harder. 💾 OLD WAY → Static Functions Remember the CTRL+SHIFT+ENTER dance? Static functions meant: - Single cell outputs that limited your analysis - Array formulas that required special keystrokes - Range references that broke easily 💻 NEW WAY → Dynamic Array Functions Welcome to the future: ✅ Arrays spill automatically where needed ✅ UNIQUE, FILTER, and SORT work instantly ✅ Dynamic ranges adjust automatically ✅ Formulas actually make sense when you read them Calculations that used to take hours now happen in seconds. 💾 OLD WAY → Basic Support Getting help meant: - Scrolling through endless forum posts - Testing solutions that might not work - Dealing with outdated answers - Wasting time on trial and error 💻 NEW WAY → Excel Copilot AI joins your team: ✅ Understands what you're trying to do ✅ Writes formulas in plain language ✅ Suggests improvements as you work ✅ Solves problems in real-time It's like having an Excel expert on speed dial. === Which modern Excel feature surprised you? What's still on your learning list? Drop your thoughts below 👇

  • View profile for Arpit Bhayani
    Arpit Bhayani Arpit Bhayani is an Influencer
    293,611 followers

    GitHub used to have a monolithic database handling 950,000 txn/sec 🤯 But this database became a pain as they scaled. So, a natural fix was to break this database into smaller ones, but this turned out to be more complex than they thought. I just published a video dissecting their entire process and the exact steps they took to do it without any downtime while making sure none of the existing functionalities break. Give it a watch - youtu.be/Tq1fif3rcnQ Going through this video, you will understand how databases are migrated without downtime, how to approach such a complex problem one step at a time, and you will get a practical look into real-world database scaling challenges and solutions.

  • View profile for Panagiotis Kriaris
    Panagiotis Kriaris Panagiotis Kriaris is an Influencer

    FinTech | Payments | Banking | Advisor, Founder, Editor

    165,766 followers

    Banks can’t afford to miss this one. Core banking modernization has long dominated banks’ boardrooms. AI is now flipping the script as it introduces - for the first time in decades - a new operating logic. Here are their options. For years, core transformation has been treated as a slow, multi-year journey - a balancing act between stability and change. Most banks took a cautious, incremental approach: patch the legacy, layer on digital, expose some APIs, and keep moving. It was less about reinvention - and more about risk containment. But this time is different. AI isn’t just another tool to plug into aging infrastructure. It demands - and enables - a fundamentally new way of operating. It runs on real-time data, modular architecture, and continuous orchestration. Traditional core systems weren’t designed for this. They operate in batch cycles, live in product silos, and require months to adapt to even small changes. Simply adding AI into that environment will not work. 𝗕𝗮𝗻𝗸𝘀 𝗻𝗼𝘄 𝗳𝗮𝗰𝗲 𝘁𝗵𝗿𝗲𝗲 𝗯𝗿𝗼𝗮𝗱 𝗰𝗵𝗼𝗶𝗰𝗲𝘀: 1. Keep modernizing around legacy cores - slow, expensive, and increasingly misaligned with business needs. 2. Attempt a full core replacement - a bold move that requires significant investment and careful execution. 3.  Rethink the role of the core entirely - and start decoupling intelligence, orchestration, and engagement from legacy constraints. That third path is fast gaining traction. Backbase is one of the leading examples. They’ve recently launched the world’s first AI-powered Banking Platform, built not as a replacement for the core, but as a new operating system around it. Here’s the set-up: 1. Unified Digital Banking Fabric – orchestrates onboarding, servicing, lending, investing, and more, across all channels. 2. Intelligence Fabric – embeds AI services, agentic automation, predictive models, and responsible AI governance into every layer. 3. Integration Fabric – connects seamlessly to existing cores, CRMs, fintechs, and third-party systems with 60+ pre-built connectors. 4. Composable Architecture – modular, cloud-native, and flexible – banks pick what they need, when they need it, without disruption. 5. Core-Agnostic: It wraps around legacy systems, accelerating transformation without the risks and costs of core replacement. This isn’t a full core replacement - it’s a strategic rearchitecture. It allows banks to keep what’s stable in the back, while radically upgrading the front and middle with intelligence, flexibility, and speed. It's a modular, composable approach that reflects the reality most banks face: evolve without breaking everything. AI won’t wait for five-year transformation roadmaps. It’s already reshaping the game - and the banks that win will be the ones building real capabilities, redesigning processes, and deploying at speed.   Opinions: my own, Graphic source: Backbase 𝐒𝐮𝐛𝐬𝐜𝐫𝐢𝐛𝐞 𝐭𝐨 𝐦𝐲 𝐧𝐞𝐰𝐬𝐥𝐞𝐭𝐭𝐞𝐫: https://lnkd.in/dkqhnxdg 

  • View profile for Dr. Georg Angenendt

    Battery Science & Commercial Strategy | Building Category Leader | Co-Founder ACCURE

    13,838 followers

    🔋 Why No Two Battery Cells Are Truly Alike: Insights into Aging Variability Battery packs are made up of multiple individual cells, typically connected in series, which means that the weakest cell defines the overall pack performance. What’s often overlooked is that all cells age individually, leading to a spread in performance over time. Just like humans, batteries are subject to both their “genetic” makeup and environmental conditions. While local operating conditions like temperature, state of charge, and load contribute to aging differences, subtle variations from the manufacturing process itself also play a key role. Even cells from the same mass production line can exhibit noticeable differences in aging trends. In a study conducted by Baumhöfer et al., 48 mass-produced UR18650E cylindrical cells were aged under identical laboratory conditions. Initially, the performance of these cells was closely aligned. However, after several hundred cycles, the aging spread became significant, with the best-performing cells outlasting the worst by a considerable margin. This experiment shows that the aging spread cannot be ignored—real-world conditions would likely make these differences even more pronounced. In series-connected battery packs, this effect is further amplified. The weakest cell, in terms of performance, sets the limit for the entire pack, potentially reducing overall capacity and cycle life. 🔋 Why does this matter? Understanding these variances—rooted both in production and operation—is essential for optimizing battery performance in real-world applications like electric vehicles and energy storage systems. By recognizing and managing these differences, manufacturers and operators can extend battery life and improve performance. 📢 Have you experienced differences in battery cells in your applications? Share your insights in the comments! 📚 Learn More: - Discover how weak cells impact Battery Energy Storage System (BESS) capacity in our latest whitepaper: [https://lnkd.in/e-pNWRMr] - Link to the research paper is in the comments.

  • View profile for Shubham Srivastava

    Principal Data Engineer @ Microsoft CoreAI | ex-Amazon | Data Engineering

    74,044 followers

    A Senior Data Engineer candidate was asked to design an incremental ingestion pipeline during his interview at Google. Another candidate in a different loop at Facebook got the same prompt. CDC pipelines look simple until you add one layer of reality: – Add late arriving updates? Now you need watermarks, reprocessing windows, and correctness guarantees. – Add duplicates and retries? Now idempotency becomes the whole game. – Add schema changes? Now your pipeline breaks at 2 AM unless you plan compatibility. – Add backfills? Now you are doing surgery on live tables without double counting. – Add merge cost? Now your “incremental” job is slower than a full reload. Here’s my checklist of 15 things you must get right when building incremental ingestion with CDC: 1. Start with the business contract → Define what “correct” means: latest state per entity, full history, or both. This single decision changes your table design, merges, and backfills. 2. Choose the right ingestion model: snapshot + CDC vs pure CDC → Snapshot + CDC is safest for bootstrapping and recovery. Pure CDC is leaner but brittle if you miss events. 3. Pick a stable primary key strategy → If your upstream keys are messy, create a durable surrogate key. Your entire dedupe and merge logic depends on this. 4. Capture an ordering signal you can trust → Use a reliable change version: log sequence number, commit timestamp, or monotonically increasing version. Avoid “updated_at” unless you fully trust the source. 5. Design for idempotency from day one  → Assume every event can arrive twice. Your writes must be safe to re-run without changing results. 6. Handle deletes explicitly → CDC isn’t just inserts and updates. Support tombstones or delete flags and define how downstream tables interpret them. 7. Preserve raw events before you transform → Land the raw change feed in a bronze layer. If downstream logic is wrong, raw becomes your rewind button. 8. Build a dedupe rule that survives retries and replays → Dedupe by (primary_key + change_version) or (primary_key + event_id). If event_id is missing, generate one deterministically from the payload plus version. 9. Use watermarks, but never trust them blindly → Watermark = “I have processed up to here.” Still keep a safety lookback window because late data is guaranteed in production. 10. Implement a reprocessing window for late arrivals  → Recompute the last N hours or days on every run based on observed lateness. This is the simplest way to get correctness without constant firefighting. 11. Plan schema evolution with compatibility rules → Decide: backward compatible only, or allow breaking changes with a controlled rollout. Use versioned schemas and block unsafe changes automatically. (Continued in comments.)

  • View profile for Sanjjeev K Singh

    CEO @ ASAR Digital | SAP Transformation Advisor | Author | Speaker

    27,561 followers

    “We need to move off SAP ECC, but we can’t justify the cost.” My CEO friend from Harvard Business School told me this last week. And she’s not wrong. Migration quotes often look scary because many SAP projects: ⚠️ Take years with runaway budgets. ⚠️ Have consultants billing endlessly without clear outcomes. ⚠️ Burn out teams testing the same processes on repeat. But why do SAP migrations become so expensive? Here’s the truth: Most SAP migrations fail the moment they’re scoped. 🛑 They scope everything instead of what matters. ➡️ Every custom report, even if no one uses it. ➡️ Every process variant, even if it’s an edge case. ➡️ Every piece of dirty data, without cleaning it first. This isn’t transformation. It’s expensive duplication. A smart SAP migration is different. It’s a business simplification project disguised as a technical upgrade. If you want to control your migration costs, here’s how: ✅ 1️⃣ Migrate only what you need. Your ECC likely has 20 years of custom code, unused reports, and workarounds that no longer serve you. S/4HANA is your chance to reset, not replicate. ✅ 2️⃣ Fix your data before you migrate. Dirty data multiplies your testing cycles and post-go-live headaches. Good data shrinks timelines, reduces consultant hours, and improves user trust. ✅ 3️⃣ Prioritize the 20% that runs 80% of your business. You don’t need to perfect every exception on day one. Get your core revenue-driving processes live, then iterate. ✅ 4️⃣ Pick a partner who says ‘no’. You need a partner who challenges scope bloat, not one who says yes to everything to grow billable hours. 🚩 Here’s what most never calculate: the cost of staying stuck. – The revenue lost because quotes take days, not hours. – The manual reconciliations your team does every month. – The friction your customers feel because your processes can’t keep up. You’re already paying a hidden cost every day you stay on ECC. You just don’t see the invoice. The difference between an expensive SAP migration and a smart one isn’t technology. It’s strategy. If migration costs are holding you back, maybe it’s time to ask: “Are we planning a migration, or are we copying our problems into a new system?” How are you thinking about controlling cost when you move off SAP ECC? #SAP #S4HANA #SAPMigration #DigitalTransformation #Leadership #CIO #CEO #EnterpriseIT #CloudERP #BusinessTransformation #SAPCommunity #ASARDigital #ERP

  • View profile for Raul Junco

    Simplifying System Design

    147,857 followers

    This question triggered my PTSD. "50M row users table. Add is_verified BOOLEAN DEFAULT FALSE. Takes 6 hours and locks the table. How do you avoid downtime?" WHY can this lock your table? Older versions of nearly every major RDBMS rewrite EVERY row on disk when you add a column with a DEFAULT. 50M rows = massive I/O. Table locked the entire time. No reads. No writes. App is dead. Modern databases improved this a lot. - PostgreSQL can apply defaults lazily (no rewrite) - MySQL 8+ can use ALGORITHM=INSTANT In those cases, this operation can be near-instant But unless you know your exact version and behavior... That’s a risky bet in production. The safe, boring, works-everywhere approach: 1. Add the column as NULL (no default, no rewrite, instant): ALTER TABLE users ADD COLUMN is_verified BOOLEAN; 2. Set the default for future inserts: ALTER TABLE users ALTER COLUMN is_verified SET DEFAULT FALSE; 3. Backfill in small batches: UPDATE users SET is_verified = FALSE WHERE id IN ( SELECT id FROM users WHERE is_verified IS NULL LIMIT 10000 ); Loop it. 10k-50k rows per batch. Sleep between batches to let replication breathe. Schema migrations on large tables are deployment events. Yes, modern databases can make this instant. But production failures happen when you assume they will. Never run ALTER TABLE on a 50M row table without a plan.

  • View profile for Pooja Jain

    Storyteller | Data Architect | Building Scalable Data & AI Foundations for Enterprise Performance | Linkedin Top Voice 2025,2024 | Open to collaboration

    198,376 followers

    𝗗𝗮𝘁𝗮 𝗤𝘂𝗮𝗹𝗶𝘁𝘆 𝗶𝘀𝗻'𝘁 𝗮 𝘀𝗶𝗻𝗴𝗹𝗲 𝗰𝗵𝗲𝗰𝗸 -it's a continuous contract enforced across the various data layers to avoid breakage. Think about it. Planes don’t just fall out of the sky when they land. Crashes happen when people miss the little signals that get brushed off or ignored. Same thing with data. Bad data doesn’t shout; it just drifts quietly—until your decisions hit the ground. When you bake quality checks into every layer and, actually use observability tools, You end up with data pipelines that hold up. Even when things get messy. That’s how you get data people can trust. Why does this matters? Bad data costs money → Failed ML models, wrong decisions. Good monitoring catches 90% of issues automatically. → Raw Materials (Ingestion)  • Inspect at the dock before accepting delivery.  • Check schemas match expectations. Validate formats are correct.  • Monitor stream lag and file completeness. Catch bad data early.  • Cost of fixing? Minimal here, expensive later.  • Spot problems as close to the source as you can. → Storage (Raw Layer)  • Verify inventory matches what you ordered.  • Confirm row counts and volumes look normal.  • Detect anomalies: sudden spikes signal upstream issues.  • Track metadata: schema changes, data freshness, partition balance.  • Raw data is your backup plan when things go sideways. → Processing (Transformation)  • Quality control during assembly is critical.  • Validate business rules during transformations. Test derived calculations.  • Check for data loss in joins. Monitor deduplication effectiveness.  • Statistical profiling reveals outliers and distribution shifts.  • Most data disasters start right here. → Packaging (Cleansed Data)  • Final inspection before shipping to warehouse.  • Ensure master data consistency across all sources.  • Validate privacy rules: PII masked, anonymization works.  • Verify referential integrity and temporal logic.  • Clean doesn’t always mean correct. Keep checking. → Distribution (Published Data)  • Quality assurance for customer-facing products.  • Check SLAs: freshness, availability, schema contracts met.  • Monitor aggregation accuracy in data marts.  • ML models: detect feature drift, prediction degradation.  • Dashboards: validate calculations match source data.  • Once data is published, you’re on the hook. → Cross-Cutting Layers (Force Multipliers)  • Metadata: rules, lineage, ownership, quality scores  • Monitoring: freshness, volume, anomalies, downtime  • Orchestration: dependencies, retries, SLAs  • Logs: failures, patterns, early warning signs Honestly, logs are gold. Don’t sleep on them. What's your job? Design checkpoints, not firefight data incidents. Quality is built in, not inspected in. Pipelines just 𝗺𝗼𝘃𝗲 data. Quality 𝗽𝗿𝗼𝘁𝗲𝗰𝘁𝘀 your decisions. Image Credits: Piotr Czarnas 𝘌𝘷𝘦𝘳𝘺 𝘭𝘢𝘺𝘦𝘳 𝘯𝘦𝘦𝘥𝘴 𝘪𝘯𝘴𝘱𝘦𝘤𝘵𝘪𝘰𝘯.  𝘚𝘬𝘪𝘱 𝘰𝘯𝘦, 𝘳𝘪𝘴𝘬 𝘦𝘷𝘦𝘳𝘺𝘵𝘩𝘪𝘯𝘨 𝘥𝘰𝘸𝘯𝘴𝘵𝘳𝘦𝘢𝘮.

  • View profile for Dr. Brindha Jeyaraman

    Founder & CEO, Aethryx | Fractional Leader in Enterprise AI Engineering, Ops & Governance | Doctorate in Temporal Knowledge Graphs | Architecting Production-Grade AI | Ex-Google, MAS, A*STAR | Top 50 Asia Women in Tech

    21,391 followers

    Excited to share my latest dive into the intersection of high-speed data and financial regulation! As digital assets and tokenized securities gain momentum, the critical question is: How do we maintain an unquestionable, tamper-proof audit trail at massive scale? Traditional databases often fall short. My new article explores how Apache Kafka's core architecture, the immutable commit log, serves as the ideal compliance layer for regulated asset transfers. I cover: 1. The power of immutability for audit-readiness. 2. Using Schema Registry to enforce structured compliance events. 3. Enabling real-time AML/KYC checks using stream processing. 4. Strategies for long-term, WORM (Write Once, Read Many) archival. If you are building infrastructure for Fintech, Digital Assets, Trading Systems, or are focused on #RegTech, you need to see how Kafka can move compliance from an "afterthought" to a real-time capability. https://lnkd.in/g_G3myVH #Kafka #DigitalAssets #Fintech #Compliance #RegTech #StreamingData #Auditability

  • View profile for Animesh Kumar

    CTO, DataOS: Get AI-ready 80% Faster

    20,184 followers

    Bronze, Silver, Gold is a solid pattern. The concern is what we put at the end of it. Veronika Heimsbakk's article on Modern Data 101 has become a quick favourite, addressing this right at the source: the medallion transformation. A Gold layer full of clean, well-modelled Parquet tables still has a fundamental limitation: relationships are outside the data. They exist in join logic, dbt models, or SQL scripts that someone wrote eighteen months ago, which break when a schema changes upstream. When a business user asks, "show me everything we know about this customer across all our systems," the answer is: go talk to the analyst who knows which tables to join. 𝐓𝐡𝐞 𝐜𝐨𝐫𝐞 𝐢𝐬𝐬𝐮𝐞 𝐢𝐬 𝐡𝐨𝐰 𝐰𝐞 𝐢𝐝𝐞𝐧𝐭𝐢𝐟𝐲 𝐭𝐡𝐢𝐧𝐠𝐬.  Every record in most data platforms has an ID that means something only inside one system. A customer_id in the CRM, an account_id in billing, a client_ref in the external registry, all pointing at the same entity with no native way to traverse those relationships without custom join logic stitching them together. We've been building pipelines that transform data structurally for decades, while the semantic layer (what things actually mean and how they relate) gets pushed off to the application layer or left undocumented entirely. 𝐆𝐥𝐨𝐛𝐚𝐥 𝐔𝐧𝐢𝐪𝐮𝐞 𝐈𝐑𝐈𝐬 What changes when you give every entity a stable, globally unique IRI in your Silver layer isn't just metadata hygiene. It's the difference between data that describes things and data that connects things. That IRI becomes a join key that works across all your systems without you having to write the join. Map your Silver DataFrames against a shared ontology in your Gold layer, publish as RDF, and the relationships are in the data itself. The SPARQL queries that become possible at that point are a different category of question entirely. Consider impact analysis, entity resolution, and provenance all in one query: "find every entity related to financial compliance that touches this customer record, trace back to source, and tell me which systems contribute what." 𝐇𝐨𝐰 𝐃𝐂𝐀𝐓 𝐟𝐢𝐭𝐬 𝐫𝐞𝐚𝐥𝐥𝐲 𝐰𝐞𝐥𝐥 𝐡𝐞𝐫𝐞 Instead of designing a custom schema to describe your datasets, you get a W3C standard ontology built for exactly this purpose. And because your catalog metadata is RDF and your data is RDF, they are in the same graph. The catalog stops being a separate application pointing at your data and becomes part of the knowledge graph itself. Let's not ask, "Should we use a knowledge graph?" More fundamentally, ask, "What does fully enriched, connected, query-ready data look like? 🔖 𝐅𝐮𝐥𝐥 𝐭𝐞𝐜𝐡𝐧𝐢𝐜𝐚𝐥 𝐰𝐚𝐥𝐤𝐭𝐡𝐫𝐨𝐮𝐠𝐡: https://lnkd.in/gSbC_Vxt #DataArchitecture #KnowledgeGraphs #SemanticWeb #DataEngineering #Medallion

Explore categories