Showing posts with label considering. Show all posts
Showing posts with label considering. Show all posts

Thursday, March 22, 2012

(Design) Production db used as the data warehouse?

We're designing our first bi suite and we're considering not using a data warehouse at all but connecting directly to the production database via the DS. I have a nagging feeling that inherently this is not a good idea but would welcome pros and cons.

Pros
- real time data updates as fast as we can Process
- no need for ETL
- we can write named queries for use in the DSV to satisfy our data requirements

Cons
- performance hit to non BI users when BI users report if using ROLAP partitions
- potential table locking during Processing, again affecting non BI users

My gut feel is that instead we should be keeping a synchronized copy of Production as a DW to report off, but I will need stronger Cons that those above to convince my boss.

Be gentle please - my first post and I completed my fist SSAS course only yesterday. Thanks.

Are your production db home made or purchased solution from sombody?

Is the data quality 100% so you don't need data clearing and validation?

Are there any updates/upgrades of production db structure?

Are you sure, that you production db is only source of your data warehouse and there will be no another sources in next years?

How large is the volume of your production db?

How large it will be in 3-5 years?

Do you plan to archive some old transaction data from your production db?

|||

Thanks Vlad, some good points here:

Are your production db home made or purchased solution from sombody?

- The production database is also ours.


Is the data quality 100% so you don't need data clearing and validation?

- Yes, and any changes that need to be made through the interfaces would be made in the source (production) database (which would then flow through).


Are there any updates/upgrades of production db structure?

Cheers I hadn't considered this - if we change the DDL of the production database during version upgrades then this will potentially have an affect on the DSV's.


Are you sure, that you production db is only source of your data warehouse and there will be no another sources in next years?

At this stage yes - although if it's very successful I could see clients wanting to report on other external sources like their upstream ERP's. I'm not sure how having the production database as the data warehouse could relate to this though...

How large is the volume of your production db?

Our largest production databases are about 5GB.


How large it will be in 3-5 years?

Hmm, hard to say but would guess 10GB?


Do you plan to archive some old transaction data from your production db?

Very good point - for customers who have been using the software for some time and have a large amount of history they may want to use the DW as an archive and remove data from production - the approach I outlined about does not allow this easily (we would have to write a number of date queries restricting the data.


Considering the lukewarm response I got to this thread perhaps what is being proposed is not such a big deal/poor option - if I consider the cons above it sounds like we may be able to implement it in such a way. I must admit I still have a bad get feel about it but can't really pin down exactly why...


Has anyone seen cubes based directly off their production database (i.e. not a copy of it) in a real world environment? Would love some more feedback...

|||

Hi,

I have seen many SSAS solutions, and a some of them direct connected to production db. but no one of direct connected was ERP db.

I depends on your businnes. Is you production db a ERP like db, or what else? If ERP like, then you must have DWH, if you what to sleep relaxed :-) another way you get enough headache

|||It's an OLTP database but not high volume, only dozens of transactions an hour. But we're thinking of only processing nightly, and setting an expectation with the users that this will be the case.

Friday, March 16, 2012

"Schemas" in same database vs multple Sql Instances

We have a DataWarehouse project. We are considering using table names (in the same database) like this:

Acct.Payroll

...

Claims.Payroll

.....

Shipping.Payroll

...

where the "Schema" differentiates the "Payroll" tables. The tables have similar functions but are quite different internally. We expect these table to grow significantly over time.

Question: From a performance\ maximum database capacities, admin perspective what the pros/cons of each approach:

1. Schema approach

2. Multiple instances of sql (up to 16) on a single server

3.Seperate databases on the same server

TIA,

barkingdog

First, having different schemas is the simplest and "fastest" approach.

Second, would be seperate databases. There is very little overhead in accessing another database on the same server.

WAY, WAY down the list, about 412 on the performance scale, would be using seperate instances. Connections are very cpu and memory intensive as well as running the entire SQL server instance, etc. I would not recommend it for this application.

I would highly suggest using file groups and many hard drives to spread the usage in one database with multiple schemas.|||

If the assumed parts of you logical design are supposed to have some logical consistency between them, you would most probably need to use transactions that span over multiple functional partitions in your application. If so, then:

1. With multiple instances/databases you won't be able to use declarative referential integrity between the tables residing on different instances/databases.

2. With multiple instances you will be paying a much higher price of the full-blown 2-phase distribution transaction commit protocol for the transactions that span multiple instances. Although for multiple database the engine also uses a variation of a 2-phace commit protocol, but in fact it's very lightweight and does not actually incur a noticeable performance overhead.

3. With multiple instances/databases if the 2-phase coordinator becomes unavailable all the participating transaction on the other instances/databases that have prepared but have not receive the commit message will not able to make any progress and will need to wait until the coordinator comes back online. Quite often resolving situations like this requires manual intervention.

4. With multiple instances you won't have automatic deadlock detection - it might become a major administrative headache because you would need to 'detect' and 'resolve' the deadlocks manually.

5. With multiple instances/databases a sound backup strategy is more complicated to design and to implement.

Probably the only real benefit (besides somewhat higher availability if the setup is right) that one could get from going multi-instance is when a single machine, however powerful it is, is unable to handle the workload. If so then you really need a scale-out solution and in this case the disadvantages don't matter and you will have to pay the price.

Assuming that you properly partition your tables and/or design the placement of the tables in filegroups, the best possible option would be to use multiple schemas in the same database. Unless, of course, a single machine is unable to handle your workload.

Another possible exception is the case when certain database level options are not applicable to all partitions of your application, For instance – the Payroll would benefit from the row-level versioning, but the Claims for some reason wouldn’t. If you don’t want to pay the versioning overhead for the Claims then you would probably want to move it to a separate database on the same instance.

|||

From a raw performance perspective on a single database server, using schemas would be best. It also would be the easiest to develop and administer.

Using multiple databases in the same instance would perform fine, but would be a little more complicated, since your applications would need different connection strings. Backups/restores would be quicker.

With multiple instances on a single database server, the available RAM will be divided between all of the instances. You will also have to deal with different connection strings.

If you find that your single database server cannot handle the workload in the future, and you don't want to scale up, you would want to have these three tables on separate databases on separate servers to scale out instead.