mercredi 11 juillet 2012

Upgrading mid-2009 MacBook Pro 13"

Editorial: I have just came back to this blog after a long period of absence due to some major life events (moving out of Australia) and I found this post still a draft, from approximately August 2011; since I did spend the time writing it, I might as well post it. So here it is folks, the first post since I'm back is a post from the past.

I've just come back from France where I celebrated my birthday a few weeks ago.
Amongst my presents was a Western Digital 2.5" Scorpio Blue 1TB hard drive and a G.Skill DDR3 SO-DIMM 8GB (2x4GB) kit for my MacBook Pro. (yeah I'm a nerd :P)
So of course I had to document the entire procedure and take pictures of everything !
The upgrade was not all that smooth (I blame Apple, see next post) but I've got it all (almost) working.

Now of course doing these kind of modifications to your computer yourself will certainly void your warranty, may cause irreversible damage to your computer etc... do it at your own risk !

OK, Let's start by taking this baby apart shall we ?
All we need is a small philips screwdriver and a small torn screwdriver (for the hard drive) as seen on the next picture:


Don't forget to use your anti-static strap if you have one, otherwise don't forget to touch some metal part of your computer to discharge yourself of any static you may be carrying.
Failure to do so may result in damage to your computer, especially to the RAM which is very sensitive to static. I attached my anti-static strap to some metal par at the back of my desktop computer case.


Next, lie the MBP upside down (face with the Apple logo down) on some soft surface (to avoid scratching it !) and unscrew the 10 screws holding the back plate (7 shorts, 3 longs).
To remember where each went I placed them in the order I took them out, see the following pic:


Once the screws are off, the plate should come off easily; put it aside.
I'll start by upgrading the HDD.
I've read some comments on the internet about how the new 1TB HDD are too big for the MBP (12.5mm), and it won't fit. I have to say it is true that it's quite big compared to the original one (see pic):


However it fits perfectly, and I have had no trouble at all installing it.
So start by un-tightening the two screws holding the HDD bracket in place:



Then remove and store that bracket with its two screws somewhere safe.
Next, pull the transparent tape slowly to get the HDD out:


You will need to remove the connector that goes to the logic board:
To do that, just insert your nail (or something of the sort), in between the connector and the HDD, and pull slightly away from the HDD until the connector becomes loose:




There are 4 (four) mounting screws around the HDD, and if your new HDD did not come with some attached (mine did not) you will have to transfer them across. This is where you need your torx screwdriver, the position of the 4 screws are circled in red in the following picture:


Screw them back onto your new HDD and you're ready to continue:


Now store that old HDD somewhere safe; in my case I will use it later to restore my mac during the install of OS X.
Put the logic board connector on the new HDD, and place it down.


Replace the HDD mounting bracket:


And you're done for the HDD !

mercredi 4 mai 2011

Generating normals on the GPU from a height map

While reading Shader X7 I stumbled upon the article "Dynamic Terrain Rendering on GPUs Using Real-Time Tessellation" (if you don't have Shader X7 yet, get it from amazon, it's a must-read).

In the article, Natalya Tartarchuk describes a way to shader such dynamic terrain that computes the normal on-the-fly.
This is very interesting since it could allow one to dynamically modify the height map and have the normals adjust automatically.
After spending a few hours debugging the normals I finally got something visually pleasing, although I've noticed a few artifacts.
After some more testing I've noticed that there is some error involved in computing the normals on the GPU (compared to the reference algorithm on the CPU).

The first image is the reference image, the normals have been generated on the CPU using the central difference algorithm.
In order to visualize the normals, they are moved into the [0,1] range by doing Normal * 0.5 + 0.5:

Reference normals generated on the CPU

The second picture is ATI's algorithm running on the pixel shader.
I had to tweak some of their variable since it seemed to rely on "magic" (but documented) value:

ATI's algorithm

The following picture shows the error between the normals computed as GPUNormal - CPUNormal:

Divergence from CPU normals. Grey areas mean error = 0 (since 0 * 0.5 + 0.5 = 0.5).

The third picture is the central difference algorithm ported on the GPU and running in the pixel shader.

The normals generated by the central difference algorithm in the pixel shader

The resulting error

My implementation of the central difference algorithm on the GPU must be pretty bad considering the divergence. I also tried the two algorithms on the vertex shader but sadly the error was just too great (except on the flat plane).

Now the question is, is it worth it to generate the normals on the GPU ?
I would say, if you're not having a dynamic terrain, you're better off with the normals on the CPU since it leads to much greater visual quality and won't steal one of those precious sampler slot so needed on SM3 hardware.
One advantage of generating the normals on the fly could be to reduce the size of the vertex buffer (take out normal + tangent from the vertex declaration => up to 6 floats per vertex removed) but that is to be balanced with the size of the height map.

mercredi 16 mars 2011

Circumventing the C# reference to reference problem

Recently I have been working on a simple recursive descent parser in C# .Net 4.0 and I encountered an error I did not expect; the following is not valid C#:

Rule integer, factor, expression, term;

integer = +Parser.digit_p();
   
factor = integer
 | Parser.ch_p('(') + expression + Parser.ch_p(')')
 | (Parser.ch_p('-') + factor)
 | (Parser.ch_p('+') + factor);
term = factor *((Parser.ch_p('*') + factor) | (Parser.ch_p('/') + factor));
expression = term * ((Parser.ch_p('+') + factor) | (Parser.ch_p('-') + factor));

This makes sense because we cannot use the rules before they have been initialized !
In C++ we'd just take a reference or pointer to the rule until they are defined; but this does not work in C#.

Instead I had to create a simple RuleRef class as follow:
public class RuleRef
{
 private Rule ptr;

 public Rule Value
 {
  get
  {
   return ptr;
  }

  set
  {
   ptr = value;
  }
 }

 public static implicit operator Rule(RuleRef r)
 {
  return r.Value;
 }

 public RuleRef()
 {

 }

 public RuleRef(Rule value)
 {
  ptr = value;
 }

 public static RuleRef operator +(RuleRef lhs, RuleRef rhs)
 {
  return new Sequence(lhs, rhs);
 }

 public static RuleRef operator +(RuleRef lhs)
 {
  return new OneOrMore(lhs);
 }

 public static RuleRef operator |(RuleRef lhs, RuleRef rhs)
 {
  return new Or(lhs, rhs);
 }

 public static RuleRef operator &(RuleRef lhs, RuleRef rhs)
 {
  return new Sequence(lhs, rhs);
 }

 public static RuleRef operator *(RuleRef lhs, RuleRef rhs)
 {
  return new Sequence(lhs, new ZeroOrMore(rhs));
 }
}

This allow me to then write:
RuleRef integer = new RuleRef();
RuleRef factor = new RuleRef();
RuleRef expression = new RuleRef();
RuleRef term = new RuleRef();

integer = +Parser.digit_p();
   
Rule _factor = integer.Value.WithAction(DebugPrint)
 | Parser.ch_p('(') + expression + Parser.ch_p(')')
 | (Parser.ch_p('-') + factor)
 | (Parser.ch_p('+') + factor);

Rule _term = factor *((Parser.ch_p('*') + factor) | (Parser.ch_p('/') + factor));

Rule _expression = term * ((Parser.ch_p('+') + factor) | (Parser.ch_p('-') + factor));

// Resolve the references to their correct values
factor.Value = _factor;
term.Value = _term;
expression.Value = _expression;

In the next post I will detail my simple recursive descent parser entirely written in C#/Linq.

lundi 13 décembre 2010

API independent primitive topology

Here is a simple uber function for retrieving the primitive topology in a native API format from an API independent format:

enum PrimitiveTopology
{
 PT_POINTS = 0,
 PT_LINES,
 PT_LINE_STRIP,
 PT_TRIANGLES,
 PT_TRIANGLE_STRIP,
 PT_TRIANGLE_FAN,
 PT_QUADS,
 PT_QUAD_STRIP
};

enum GraphicsAPI
{
 DirectX11,
 DirectX10,
 DirectX10_1,
 DirectX9,
 OpenGL,
 OpenGLES
};

template <GraphicsAPI API>
int getPrimitiveTopology(PrimitiveTopology pt)
{
 switch (API)
 {
  case DirectX11:
  case DirectX10:
  case DirectX10_1:
   switch(pt)
   {
    case PT_POINTS:
     return D3D_PRIMITIVE_TOPOLOGY_POINTLIST;
     break;
     
    case PT_LINES:
     return D3D_PRIMITIVE_TOPOLOGY_LINELIST;
     break;
     
    case PT_LINE_STRIP:
     return D3D_PRIMITIVE_TOPOLOGY_LINESTRIP;
     break;
     
    case PT_TRIANGLES:
     return D3D_PRIMITIVE_TOPOLOGY_TRIANGLELIST;
     break;
     
    case PT_TRIANGLE_STRIP:
     return D3D_PRIMITIVE_TOPOLOGY_TRIANGLESTRIP;
     break;
     
    case PT_TRIANGLE_FAN:
     return D3D_PRIMITIVE_TOPOLOGY_UNDEFINED;
     break;
     
    default:
     assert(0 && "getPrimitiveTopology::DirectX10+ Unknown or unsupported primitive topology !");
     return -1;
     break;
   }
   break;
  case DirectX9:
   switch(pt)
   {
    case PT_POINTS:
     return D3DPT_POINTLIST;
     break;
     
    case PT_LINES:
     return D3DPT_LINELIST;
     break;
     
    case PT_LINE_STRIP:
     return D3DPT_LINESTRIP;
     break;
     
    case PT_TRIANGLES:
     return D3DPT_TRIANGLELIST;
     break;
     
    case PT_TRIANGLE_STRIP:
     return D3DPT_TRIANGLESTRIP;
     break;
     
    case PT_TRIANGLE_FAN:
     return D3DPT_TRIANGLEFAN;
     break;
     
    default:
     assert(0 && "getPrimitiveTopology::DirectX9 Unknown or unsupported primitive topology !");
     return -1;
     break;
   }
   break;
  case OpenGL:
  case OpenGLES:
   switch(pt)
   {
    case PT_POINTS:
     return GL_POINTS;
     break;
     
    case PT_LINES:
     return GL_LINES;
     break;
     
    case PT_LINE_STRIP:
     return GL_LINE_STRIP;
     break;
     
    case PT_TRIANGLES:
     return GL_TRIANGLES;
     break;
     
    case PT_TRIANGLE_STRIP:
     return GL_TRIANGLE_STRIP;
     break;
     
    case PT_TRIANGLE_FAN:
     return GL_TRIANGLE_FAN;
     break;
    default:
     assert(0 && "getPrimitiveTopology::OpenGL/ES Unknown or unsupported primitive topology !");
     return -1;
     break;
   }
   break;
  default:
   assert(0 && "getPrimitiveTopology() Unknown API !");
   return -1;
   break;
 }
}


A good compiler should optimize the dead code.
Also if you want to avoid the dependency on the graphics API headers, here are the necessary enums:


enum D3D_PRIMITIVE_TOPOLOGY
{
 D3D_PRIMITIVE_TOPOLOGY_UNDEFINED = 0,
 D3D_PRIMITIVE_TOPOLOGY_POINTLIST = 1,
 D3D_PRIMITIVE_TOPOLOGY_LINELIST = 2,
 D3D_PRIMITIVE_TOPOLOGY_LINESTRIP = 3,
 D3D_PRIMITIVE_TOPOLOGY_TRIANGLELIST = 4,
 D3D_PRIMITIVE_TOPOLOGY_TRIANGLESTRIP = 5
};

enum D3DPRIMITIVETYPE
{
 D3DPT_POINTLIST = 1,
 D3DPT_LINELIST = 2,
 D3DPT_LINESTRIP = 3,
 D3DPT_TRIANGLELIST = 4,
 D3DPT_TRIANGLESTRIP = 5,
 D3DPT_TRIANGLEFAN = 6,
 D3DPT_FORCE_DWORD = 0x7fffffff,
};

enum GLPrimitiveTopology
{
 GL_POINTS                        = 0x0000,
 GL_LINES                         = 0x0001,
 GL_LINE_LOOP                     = 0x0002,
 GL_LINE_STRIP                    = 0x0003,
 GL_TRIANGLES                     = 0x0004,
 GL_TRIANGLE_STRIP                = 0x0005,
 GL_TRIANGLE_FAN                  = 0x0006,
 GL_QUADS                         = 0x0007,
 GL_QUAD_STRIP                    = 0x0008,
 GL_POLYGON                       = 0x0009,
};

If you define the native types yourself, watch out for conflicts !

lundi 11 octobre 2010

What do programmers do when they are bored ?

Ok so I have not posted in a long time because I've been very busy with my uni stuff; so I thought I might share a nice little trick I developed quickly this morning.

Yes I was very bored and skimming through DXUT source file when I spotted those:
#define SET_ACCESSOR( x, y )       inline void Set##y( x t )   { DXUTLock l; m_state.m_##y = t; };
#define GET_ACCESSOR( x, y )       inline x Get##y()           { DXUTLock l; return m_state.m_##y; };
#define GET_SET_ACCESSOR( x, y )   SET_ACCESSOR( x, y ) GET_ACCESSOR( x, y )

#define SETP_ACCESSOR( x, y )      inline void Set##y( x* t )  { DXUTLock l; m_state.m_##y = *t; };
#define GETP_ACCESSOR( x, y )      inline x* Get##y()          { DXUTLock l; return &m_state.m_##y; };
#define GETP_SETP_ACCESSOR( x, y ) SETP_ACCESSOR( x, y ) GETP_ACCESSOR( x, y )

and I instantly wanted the same type of macros, except with all the genericity and safety of C++0x.

Setters were a walk in the park:
#define Nitro_PropertySet(name)    inline void set##name(decltype(m##name) t) { m##name = t; }

Getters not so much; I initially thought that I could use the trailing-type feature of C++0x to get an auto-magic return type, unfortunately it behaves differently and would not accept the member name for an answer.
So I used a little trick, I declare a template structure with a default template parameter (the decltype with the member name), add a typedef, inject the whole in the inline getters and there it is !

#define Nitro_PropertyGet(name)    template < typename T = decltype(m##name)> struct __##name { typedef typename T type; }; inline __##name<>::type get##name() { return m##name; }

For a member named mNumSplits, that expands to:
 template < typename T = decltype(mNumSplits)>
 struct __mNumSplits { typedef typename T type; };
 inline __mNumSplits<>::type getNumSplits() { return mNumSplits; }

Neat. And all that while eating breakfast :)

mercredi 1 septembre 2010

A horrible hack turned safe thanks to templates.

While coding some shader implementation for DirectX 11 I found myself writing the same code over and over for each shader type (Vertex, Fragment, Geometry, Hull, Domain etc...) and I figured I should make it more generic.

The only problem is that Direct3D 11 has a separate function to create each type of shader, so in order to make it generic I originally came up with a solution which I qualify as "horrible hack".

In DirectX 11 the ID3D11Device has some methods like:
HRESULT __stdcall ID3D11Device::CreateVertexShader(
    [in]   const void *pShaderBytecode,
    [in]   SIZE_T BytecodeLength,
    [in]   ID3D11ClassLinkage *pClassLinkage,
    [out]  ID3D11VertexShader **ppVertexShader);


    HRESULT __stdcall ID3D11Device::CreatePixelShader(
    [in]   const void *pShaderBytecode,
    [in]   SIZE_T BytecodeLength,
    [in]   ID3D11ClassLinkage *pClassLinkage,
    [out]  ID3D11PixelShader **ppPixelShader
    );

And ID3D11VertexShader and ID3D11PixelShader are both inheriting publicly from ID3D11DeviceChild:

ID3D11VertexShader : public ID3D11DeviceChild { ... }
    ID3D11PixelShader : public ID3D11DeviceChild { ... }

Now the trick was to define a function pointer CreateShader that has the following prototype:

typedef HRESULT (__stdcall ID3D11Device::*CreateShader)(const void *,SIZE_T, ID3D11ClassLinkage*,ID3D11DeviceChild**);

The only difference is the last argument which points to the superclass of the shader classes.
Now I only have to pass this function pointer pointing to the right method of the ID3D11Device when I create my shader.
The problem with that is that I have to pass the shader as its base type, which is supposed to work according to the rules of C++ except, however the (wise) compiler complained when I tried to do it implicitly.
That forced me to create a ID3D11DeviceChild* and assign the ID3D11*Shader to it, then pass it to the method.
In other word it look even more horrible; that's where templates come to help.

I decided to template the loadShader method, and use the shader type to select the appropriate ID3D11::Create* method.
Since function templates cannot have defaults, I created a helper structure:

template <class Shader>
struct CreateShaderHelper
{
 typedef typename HRESULT (__stdcall ID3D11Device::*FuncType)(const void *,SIZE_T, ClassLinkage*,Shader**);
};
Now I can re-write the loadShader method like this:

template <class Shader>
int loadShader(const char* name,
        const char* defines,
        const char* profile,
        typename CreateShaderHelper<Shader>::FuncType shader_creator,
        ConstantBuffer& cbuff,
        Shader*& outShader,
        ID3D11ShaderReflection*& outShaderReflect)
{
    ...
    if ((device->*shader_creator)(shader_buffer->GetBufferPointer(),shader_buffer->GetBufferSize(),&outShader) == S_OK)
    {
         ...
    }
}

lundi 30 août 2010

RenderTarget or Camera centric engine design ?

The engine I'm currently working on has a RenderTarget centric design, a lot like Ogre's.
A RenderTarget centric design means that all the rendering is done around RenderTargets:
The rendering window is a render target, the shadow maps are render targets etc...

Essentially a render target can have viewports, and each viewport has a camera.
This is good for doing things like multi-viewports views in games (like split screens), and it's easy to just add a viewport to a render target and configure a camera for it.
An advantage of this technique is that it minimizes render target changes, you set the render target, collect all viewports and cameras associated with it, then get all the visible objects and render them; rince repeat.
Obviously the downside is a lot of changes in camera matrices and viewports.

Recently I have acquired the book Game Engine gems 1, and in particular the gem by Colt McAnlis from Blizzard Entertainment caught my attention.
He describes a camera-centric engine design for multi-threaded rendering.
In his design everything is a camera and his association is done via a RenderView structure:
The RenderView struct groups the camera, frustum, render target and all the rendering commands.
This allows him to build drawing command buffers in parallel and then submit them to the API, grouped to minimize state changes.

At first I was bit against his design (after all we're all a bit reluctant to changes especially after working so long on a different design that does the job), but I realized that a camera-centric approach might facilitate dealing with things like portal-rendering which my design didn't account for at all.
That doesn't mean it's not possible to do with my current engine, just that I didn't think of that and I believe that a camera-centric design makes more sense; every camera contribute their view data to the final scene, including portal cameras.

In conclusion, RenderTarget centric design starts from a render target and move toward cameras, whereas camera-centric designs move from cameras to render targets.
Both design are good, and help structure the rendering engine, but a camera-centric design might reflect more how we think about cameras in real life; at least compared to the RenderTarget design which might seem backwards.